UPDATE: Merged the NoWarning into a real DPO for later use. Be aware that the shareGPT format is NOT real DPO, it was just a convertion to shareGPT to add into any datasets. If you want to do a REAL DPO train, use this file: toxic-dpo-NoWarning.json.
DISCLAIMER : I'M NOT THE AUTHOR OF THIS DATASET.
ALL CREDIT GO TO unalignment repo.
ORIGINAL DATASET: unalignment/toxic-dpo-v0.1
I just converted/modified the dataset! Only the accepted replies was taken for the shareGPT format!… See the full description on the dataset page:
https://huggingface.co/datasets/Undi95/toxic-dpo-v0.1-sharegpt.