RM-R1 is a training framework for Reasoning Reward Model (ReasRM) that judges two candidate answers by first thinking out loud—generating structured rubrics or reasoning traces—then emitting its preference. Compared to traditional scalar or generative reward models, RM-R1 delivers state-of-the-art performance on public RM benchmarks on average while offering fully interpretable… See the full description on the dataset page:
https://huggingface.co/datasets/gaotang/RM-R1-Reasoning-RLVR.