This dataset is designed for ORPO or DPO training.
See Uncensor any LLM with Abliteration for more information about how to use it.
This is version with raw text instead of lists of dicts as in the original version here.
It makes easier to parse in Axolotl, especially for DPO.ORPO-DPO-mix-40k-flat is a combination of the following high-quality DPO datasets:
argilla/Capybara-Preferences: highly scored chosen answers >=5 (7,424 samples)… See the full description on the dataset page:
https://huggingface.co/datasets/mlabonne/orpo-dpo-mix-40k-flat.