This dataset is a conversion of nvidia/HelpSteer2 into preference pairs based on the helpfulness score for training DPO.
HelpSteer2-DPO is also licensed under CC-BY-4.0.
In accordance with the following paper, HelpSteer2: Open-source dataset for training top-performing reward models
we converted nvidia/HelpSteer2 dataset into a preference dataset by taking the response with the higher helpfulness score as the chosen response, with the remaining response being the… See the full description on the dataset page:
https://huggingface.co/datasets/Atsunori/HelpSteer2-DPO.