A DPO alignment dataset for fine tuning open-source LLMs, taking sample preference pairs from a variety of datasets for diversity.
Prepared in the "standard" instruction, chosen, and rejected format, with a source feature indicating from which dataset the row was extracted.
Source datasets:
argilla/distilabel-math-preference-dpo
jondurbin/truthy-dpo-v0.1
argilla/distilabel-intel-orca-dpo-pairs
argilla/OpenHermes2.5-dpo-binarized-alpha