Qwen3-1.7B Process Reward Model for Math
Training target: Qwen/Qwen3-1.7B
Intended method: TRL PRMTrainer on step-level math supervision from trl-lib/math_shepherd.
Status: repository created. Training script/job launch pending due sandbox duplication rate-limit in the current session.