This contains the human preference evaluation set for Preference Proxy Evaluations.
This dataset is meant for benchmarking and evaluation, not for training.
Paper
Code
User prompts are licensed under CC-BY-4.0, and model outputs are governed by the terms of use set by the respective model providers.
@misc{frick2024evaluaterewardmodelsrlhf,
title={How to Evaluate Reward Models for RLHF},
author={Evan Frick and Tianle Li and… See the full description on the dataset page:
https://huggingface.co/datasets/lmarena-ai/PPE-Human-Preference-V1.