AI-feedback preference pairs used to align the legal/financial SLMs via DPO and RLAIF.
On-policy candidates were sampled from each SFT model and ranked by gpt-4.1-mini into
(chosen, rejected) pairs.
Schema: 125M/500M = {system, user, chosen, rejected, chosen_score, rejected_score};
Gemma… See the full description on the dataset page:
https://huggingface.co/datasets/Sudhanshu1985/slm-preference-pairs.