What Lies Beneath: A Call for Distribution-based Visual Question & Answer Datasets
Publication: JCDL 2025 Website and on arXiv
GitHub Repo: ReadingTimeMachine/LLM_VQA_JCDL2025
This is a histogram-based dataset for visual question and answer (VQA) with humans and large language/multimodal models (LMMs).
Data contains synthetically generated single-panel histograms images, data used to create histograms, bounding box data for titles, axis and tick labels, and… See the full description on the dataset page: https://huggingface.co/datasets/ReadingTimeMachine/visual_qa_histograms.