This is the filtered version of SHP dataset, which was used to train MMPO, as introduced in the paper below:
Margin Matching Preference Optimization: Enhanced Model Alignment with Granular Feedback
Kyuyoung Kim*, Ah Jeong Seo*, Hao Liu, Jinwoo Shin, Kimin Lee
In EMNLP 2024 Findings
The original SHP dataset consists of 385k collective human preferences over responses to questions/instructions in 18 different… See the full description on the dataset page:
https://huggingface.co/datasets/Ahjeong/SHP_filtered_MMPO.