Clean Alignment Dataset is a safety preference dataset for Direct Preference
Optimization (DPO) and related preference-alignment methods. Every example is a
(prompt, chosen, rejected) triple in which the chosen response is safe and
the rejected response is unsafe for the same prompt — an unambiguous,
consistently-labelled safe-vs-unsafe contrast in every single pair.
It is built by combining and re-cleaning two… See the full description on the dataset page: https://huggingface.co/datasets/etrigan5500/Clean-Alignment-Dataset.