Fixed version of laion/tulu3-sft-personas-math-sandboxes-verified.
The original dataset's tests/test.sh verifier emitted Correct answer: N / Incorrect answer: expected N, got M — a format the pass_ratio reward shaper cannot parse. The shaper silently fell back to binary reward (reward_shaping_fallback: true), collapsing all rewards to {0.0, 1.0} and providing no gradient signal for RLOO training.
Fix:… See the full description on the dataset page:
https://huggingface.co/datasets/laion/tulu3-sft-personas-math-sandboxes-verified-v2.