120K Visual Spatial Description dataset for instruction-tuning Large Language-and-Vision Assistant.
Visual Spatial Description (VSD) aims to generate texts that describe the spatial relationships between objects within images.
Traditional visual spatial relationship classification (VSRC) methods typically output the spatial relationship between two objects in an image,
often neglecting world knowledge and lacking general language… See the full description on the dataset page:
https://huggingface.co/datasets/swordli/LLaVA-VSD-120K.