A high-quality DPO (Direct Preference Optimization) dataset with intelligent refusal filtering applied to ensure all "chosen" responses comply with requests without censorship. This dataset combines synthetic generation via VellumForge2 with advanced refusal detection to create training data for models that handle mature content appropriately without unnecessary refusal behaviors.
This dataset underwent a… See the full description on the dataset page:
https://huggingface.co/datasets/rx1lora/tb00-VellumK2-Unfettered-DPO-01.