This dataset contains a collection of final-answer mathematical problems used for the ProofRank evaluation benchmark. It combines high-school level competition problems from MathArena (AIME, HMMT, Apex) and IMO-AnswerBench to evaluate large language models on proof quality metrics beyond basic correctness.
main: Contains 382 base problems used to… See the full description on the dataset page:
https://huggingface.co/datasets/Anon987281293/ProofRank.