A high-quality, leakage-free binary classification dataset for detecting prompt injection and jailbreak attacks against Large Language Models.
Zero data leakage — group-aware splitting confirmed
Balanced classes — ~60% malicious / 40% benign
Two configs — core for classical ML, full for transformers
29 attack categories including cutting-edge 2025 techniques
Severity labels, source tracking, augmentation flags on every row… See the full description on the dataset page:
https://huggingface.co/datasets/cyberec/Prompt-injection-dataset2.