UniRRM-RL: Reinforcement Learning Data for Unified Reasoning Reward Models
Overview
UniRRM-RL is the reinforcement learning (RL) dataset used in the second training stage of UniRRM, a Unified Reasoning Reward Model. It contains 32,832 samples in a hybrid format combining both pairwise (chosen/rejected) and listwise (ABCD four-choice) evaluation paradigms, covering 106 languages and multiple domains.
This dataset is introduced in the following paper, accepted at… See the full description on the dataset page: https://huggingface.co/datasets/SUSTech-NLP/UniRRM-RL.