CapMIT1003 is a dataset of captions and click-contingent image explorations collected during captioning tasks.
CapMIT1003 is based on the same stimuli from the well-known MIT1003 benchmark, for which eye-tracking data
under free-viewing conditions is available, which offers a promising opportunity to concurrently study human attention under both tasks.
capmit1003_dataset =… See the full description on the dataset page:
https://huggingface.co/datasets/azugarini/CapMIT1003.