This dataset was used to fine-tune the base models to be reference models in the paper CleanGen. The dataset contains 1800 conversations from UltraChat and 200 samples from HH-RLHF. For each harmful question from HH-RLHF, a refusal phrase, "I'm sorry, but I cannot assist with that," is added at the beginning of the response.
For more details, see the following paper:
CleanGen: Mitigating Backdoor Attacks for Generation Tasks in Large Language Models… See the full description on the dataset page:
https://huggingface.co/datasets/TaiGary/base_model_fine_tune_data_ultrachat_2k.