Training dataset for fine-tuning language models to resist prompt injection attacks. Created for the Prompt Injection Challenge - an AI security challenge where users attempt to extract a hidden flag from a chatbot.
Single-turn conversations with immediate prompt injection attempts and polite refusals. Examples include role-playing, system prompt overrides, and… See the full description on the dataset page:
https://huggingface.co/datasets/Alindstroem89/guardrail-training-dataset.