R3-Dataset-14k is a dataset we curated to train rubric reward models for R3, a series of Robust Rubric-Agnostic Reward Models.
We begin with a large pool of publicly available datasets spanning over 1 million examples, which include general chat, reasoning, and classification tasks and then enrich each example with on-the-fly rubric
generation and explanation traces. Finally, we apply filtering and refinement to produce smaller, higher-quality datasets used in… See the full description on the dataset page:
https://huggingface.co/datasets/rubricreward/R3-Dataset-14K.