SPRIGHT (SPatially RIGHT) is the first spatially focused, large scale vision-language dataset. It was built by re-captioning
∼6 million images from 4 widely-used datasets:
This repository contains the re-captioned data from CC12M and Segment Anything, while the COCO data is present here. We do not release images from LAION, as the parent images are currently private.
Below are some illustrative examples… See the full description on the dataset page:
https://huggingface.co/datasets/SPRIGHT-T2I/spright.