This dataset was developed and used for the SciVQA: Scientific Visual Question Answering Shared Task hosted at the Scholary Document Processing workshop at ACL 2025.
Competition is available on Codabench.
SciVQA is a corpus of 3000 real-world figure images extracted from English scientific publications available in the ACL Anthology and arXiv.
The figure images are collected from the two pre-existing datasets, ACL-Fig and SciGraphQA.
All figures are… See the full description on the dataset page:
https://huggingface.co/datasets/katebor/SciVQA.