This contains the MATH correctness preference evaluation set for Preference Proxy Evaluations.
The prompts are sampled from MATH.
This dataset is meant for benchmarking and evaluation, not for training.
Paper
Code
User prompts are licensed under MIT, and model outputs are governed by the terms of use set by the respective model providers.
@misc{frick2024evaluaterewardmodelsrlhf,
title={How to Evaluate Reward Models for… See the full description on the dataset page:
https://huggingface.co/datasets/lmarena-ai/PPE-MATH-Best-of-K.