Document parsing refers to the process of converting complex documents, such as PDFs and scanned images, into structured text formats like HTML and Markdown.
It is especially useful as a preprocessor for RAG systems, as it preserves key structural information from visually rich documents.
While various parsers are available on the market, there is currently no standard evaluation metric to assess their performance.
To address this gap, we… See the full description on the dataset page:
https://huggingface.co/datasets/upstage/dp-bench.