Dataset with reference completions and rewards for a specific model and reward model, ready for training with the QRPO reference codebase (
https://github.com/CLAIRE-Labo/quantile-reward-policy-optimization).
Part of the dataset collection for the paper Quantile Reward Policy Optimization: Alignment with Pointwise Regression and Exact Partition Functions (
https://arxiv.org/pdf/2507.08068).