Very few preference datasets have heldout test sets for validation of reward model accuracy results.
In this dataset, we curate the test sets from popular preference datasets into a common schema for easy loading and evaluation.
Anthropic HH (Helpful & Harmless Agent and Red Teaming), test set in full is 8552 samples
Anthropic HHH Alignment (Helpful, Honest, & Harmless), formatted from Big Bench for standalone evaluation.
Learning to summarize, downsampled from… See the full description on the dataset page:
https://huggingface.co/datasets/allenai/preference-test-sets.