slm-500m-rlaif2
500m legal SLM: base -> QA SFT -> instruction SFT -> RLAIF v2. A
Bradley-Terry reward model (warm-started from the instruction model) was
trained on ~4k AI-written preference pairs spanning closed-book QA AND
instruction-following failure modes; the policy was then improved by
best-of-N reward-weighted SFT.