Queries, qrels, and source manifests for six visual-document retrieval benchmarks:
InfoVQA, ChartQA, SlideVQA, TQA, OWID Charts, and Wikimedia Maps.
Retrieval is performed separately in each benchmark's own corpus. This repository includes
the complete corpus images for all six benchmarks as embedded bytes in sharded Parquet
files; no external image paths are required at evaluation time.
Each dataset directory contains query.parquet, qrels.txt… See the full description on the dataset page:
https://huggingface.co/datasets/hmhm1229/ConceptFormer-Eval.