A benchmark of 100 arXiv mathematics papers that contained an error later
corrected by the authors in a subsequent version, each annotated with the
location of the error. Introduced in Pseudo-Formalization for Automatic
Proof Verification to evaluate LLM proof verifiers on real research-level
mathematics.
The benchmark was first released with 35 papers and later expanded to 100.
The benchmark_version column records which… See the full description on the dataset page:
https://huggingface.co/datasets/LukeBailey181Pub/ArxivMathGradingBench.