WebMainBench is a high-precision benchmark for evaluating web main content extraction. It provides:
A 7,809-page, 100% human-annotated evaluation dataset covering 5,434 unique domains, 150 TLDs, and 46 languages.
A 545-sample subset with manually calibrated ground-truth markdown (groundtruth_content), enabling fine-grained metric evaluation across text, code, formula, and table dimensions.
A unified evaluation toolkit… See the full description on the dataset page:
https://huggingface.co/datasets/opendatalab/WebMainBench.