This dataset is a 1K subset of the Flickr30k VQA dataset, including per-word attention maps extracted with Stable Diffusion v1.4.
image: Main image associated with the question, stored in the images/ subfolder.
question: VQA question derived from Flickr30k captions.
answers: Ground-truth answers.
attention_images: A dictionary of per-word saliency images stored in the attention_images/ subfolder.
All images… See the full description on the dataset page:
https://huggingface.co/datasets/lxasqjc/flickr1k-sd-attn.