μ-MATH (Meta U-MATH) is a meta-evaluation dataset derived from the U-MATH benchmark.
It is intended to assess the ability of LLMs to judge free-form mathematical solutions. The dataset includes 1,084 labeled samples generated from 271 U-MATH tasks, covering problems of varying assessment complexity.
For fine-grained performance evaluation results, in-depth analyses and detailed discussions on behaviors and biases of LLM judges, check out our paper.
📊 U-MATH benchmark at Huggingface
🔎 μ-MATH… See the full description on the dataset page:
https://huggingface.co/datasets/toloka/mu-math.