This dataset comprises images and annotations from the original Relaion Synthetic Dataset.
Out of the 115M images, a subset of 5.1M images has been annotated with automatic methods (Image-text-to-text models).
dense_caption: A dense annotation about the image
vqa: Visual Question-Answers related to the image. JSON dictionary… See the full description on the dataset page:
https://huggingface.co/datasets/Fhrozen/relaion-synthetic.