This dataset was introduced in the paper Benchmarking Open-Source Layout Detection Models for Data Snapshot Extraction from Institutional Documents.
The source code for the benchmark and dataset extraction is available on GitHub: worldbank/ai4data.
The data-snapshot dataset is an annotated corpus designed for the evaluation and development of models for extracting data snapshots from PDF documents. A data snapshot is… See the full description on the dataset page:
https://huggingface.co/datasets/ai4data/data-snapshot.