This dataset contains harmful prompts and potentially harmful model reasoning traces. The prompts are sourced from HarmBench and StrongReject datasets, which include requests for illegal activities, harmful instructions, and other problematic content. The model reasoning traces may contain discussions of harmful methods, even when the models ultimately decides to refuse the requests.
This dataset is intended for AI safety research purposes only. Users should… See the full description on the dataset page:
https://huggingface.co/datasets/AISafety-Student/reasoning-safety-behaviours.