Project Page | Paper | Code
This is the dataset used to train mR3, a massively multilingual, rubric-agnostic reward reasoning model.
The mR3 training dataset contains 100,000 high-quality samples curated from an initial pool of 4 million samples across 125 languages. It is designed to train reward models that can provide reasoning traces in both English and non-English settings, covering 72… See the full description on the dataset page:
https://huggingface.co/datasets/rubricreward/mR3-Dataset-100K-EasyToHard.