A benchmark dataset of 206 programmatically generated PDF documents with deterministic ground truth markdown, designed to evaluate OCR and Vision-Language Model (VLM) performance on tabular document parsing.
Edge-case tables
43
Stress-test scenarios: empty cells, merged headers… See the full description on the dataset page:
https://huggingface.co/datasets/roma2025/synthetic-table-bench.