A large-scale dataset designed to teach LLMs how to safely refuse jailbreak attempts, prompt injections, and policy-violating requests.
This dataset contains 30GB of (category, prompt, response) triplets pairing simulated adversarial prompts with safe, helpful refusals. The data is non-operational and does not contain real exploits or harmful instructions.
Column
Type
Description… See the full description on the dataset page:
https://huggingface.co/datasets/Gugu8/Jailbreak-Refusal.