ArXiv AI Figure-Caption Dataset
This dataset contains figures and their captions extracted from recent AI-related papers on arXiv.
Total figures: 17094
Source: arXiv papers from 2023.12 onwards
Categories: CS.AI, CS.LG, CS.CV, CS.CL, CS.NE, STAT.ML
Each entry contains:
image: URL of the figure image
text: Caption text
paper_id: arXiv paper ID
figure_idx: Index of the figure in the paper