Rasterized page images for the arXiv AI/ML OCR benchmark corpus. Pages are
rendered at 144 DPI, encoded as WebP (quality=85,
method=6), and packed into parquet shards with the Hugging
Face Image feature so datasets.load_dataset decodes them automatically.
Source PDFs: the curated page-bounded dataset obswork/arxiv-ai-ml-100k-pages.
100,056 pages across 4,866 papers
Categories:
cs.AI: 25,015 pages
cs.CV: 25,011 pages
cs.LG: 25,029 pages
stat.ML:… See the full description on the dataset page:
https://huggingface.co/datasets/obswork/arxiv-ocr-benchmark-corpus.