Reformatted from nvidia/HelpSteer dataset.
The LION-series are trained using an empirically optimized pipeline that consists of three stages: SFT, DPO, and online preference learning (online DPO). We find simple techniques such as sequence packing, loss masking in SFT, increasing the preference dataset size in DPO, and online DPO training can significantly improve the performance of language models. Our best models (the LION-series) exceed the… See the full description on the dataset page:
https://huggingface.co/datasets/Columbia-NLP/DPO-HelpSteer.