This dataset is a mirror of the table dimension of llamaindex/ParseBench, packaged together with a curated Financial Split that we built for evaluating OCR systems on insurance and financial filings.
It contains,
All 503 PDFs of the ParseBench table track.
table.jsonl, the original ground truth (one HTML table per page, plus easy or hard difficulty tag).
financial_split/, our 151 page financial slice plus the 117 dropped non financial… See the full description on the dataset page:
https://huggingface.co/datasets/roma2025/parsebench-table-track.