A 50-document subset of LongExtractBench, a benchmark for schema-driven structured
extraction from long, table-heavy PDFs. Each example pairs a real-world document with a
target JSON Schema and a human-reconciled ground-truth extraction, so any extraction
system — a document-AI platform or a general LLM — can be scored field-by-field on an
identical workload.
50 documents, one folder per example (~358 MB total).
Each… See the full description on the dataset page:
https://huggingface.co/datasets/micro1-inc/longextract-bench-50.