The OpenSpaces dataset is created using VQASynth to synthesize spatialVQA data using images from the first 30K rows
of the localized narratives split of the cauldron.
Compared to the related dataset used to train SpaceLLaVA,
the OpenSpaces emphasizes greater diversity in the image distribution instead of focusing on warehouse scenes.
The following chart shows the distribution of images over tags labeled by CLIP embedding similarity:
The OpenSpaces dataset also includes… See the full description on the dataset page:
https://huggingface.co/datasets/remyxai/OpenSpaces.