This dataset is designed for ORPO or DPO training.
See Fine-tune Llama 3 with ORPO for more information about how to use it.
It is a combination of the following high-quality DPO datasets:
argilla/Capybara-Preferences: highly scored chosen answers >=5 (7,424 samples)argilla/distilabel-intel-orca-dpo-pairs: highly scored chosen answers >=9, not in GSM8K (2,299 samples)
argilla/ultrafeedback-binarized-preferences-cleaned: highly scored chosen answers >=5 (22… See the full description on the dataset page:
https://huggingface.co/datasets/mlabonne/orpo-dpo-mix-40k.