Nowadays, the possibility to generate unsafe content is fundamental. Unsafe content can be used to train/evaluate moderation models,
evaluate the safety of LLM and align LLMs to safety policies.
For example, in the Constitutional Classifier paper by Anthropic, they have used an internal model without harmlessness optimization to generate unsafe content and train
moderation models using that. Unfortunatelly, it is not always easy to generate this kind of… See the full description on the dataset page:
https://huggingface.co/datasets/fedric95/T2TSyntheticSafetyBench-v2.