slm-500m-dpo2
500m legal SLM: base -> QA SFT -> instruction SFT -> DPO v2 (beta=0.1) on ~4k AI-feedback preference pairs spanning closed-book QA AND instruction-following failure modes, prompts held out of both SFT sets. Reference = the instruction-tuned model.