This contains the MMLU-Pro correctness preference evaluation set for Preference Proxy Evaluations.
The prompts are sampled from MMLU-Pro.
This dataset is meant for benchmarking and evaluation, not for training.
Paper
Code
User prompts are licensed under MIT, and model outputs are governed by the terms of use set by the respective model providers.
@misc{frick2024evaluaterewardmodelsrlhf,
title={How to Evaluate Reward… See the full description on the dataset page:
https://huggingface.co/datasets/lmarena-ai/PPE-MMLU-Pro-Best-of-K.