RLAIF (Reinforcement Learning from AI Feedback) alignment of Sudhanshu1985/slm-125m-sft.
A scalar reward model (Bradley-Terry on gpt-4.1-mini preference pairs) scores on-policy
samples; the policy is optimized with RLOO policy gradient + KL to the frozen SFT.
Mean reward -0.158 -> 0.155. This is the RL-based counterpart to Sudhanshu1985/slm-125m-dpo.
Uses the SFT chat template: <|bos|><|system|>SYS<|user|>USER<|assistant|>ANSWER<|eos|>.