Controlled experiment testing whether scalar reward models have measurable per-axis blindness on IFEval multi-constraint prompts, and whether this blindness predicts behavior under best-of-N selection.
Open Google Colab
Change runtime to A100 GPU (Runtime → Change runtime type → A100)
Set your HF token in Colab Secrets (🔑 icon on the left sidebar): add HF_TOKEN with your token
Run this in a cell: