RLAIF (DPO (Direct Preference Optimization)) refinement of sudhisrk1982/slm-125m-instruct,
trained on 500 AI-generated preference triplets (chosen from
Gemini 2.5 Flash-Lite, rejected sampled from the SFT model itself, judged by
Gemini 2.5 Flash).
Result: coherent improvement. Answers are more complete and on-topic than the SFT baseline. DPO was the most stable of the three methods (val loss 0.30, preference accuracy 1.00).