A 23,699-row binary text-classification corpus for detecting prompt injection,
jailbreak attempts, and indirect attacks against LLM-based systems, paired with
hard adversarial negatives drawn from real human text. Every row carries its
source, license, language, and full provenance.
Rows: 23,699 (train 16,888 · validation 3,301 · test 3,510)
Labels: ataque (attack) 12,317 · benigno (benign) 11,382
Languages: English plus 8 translation… See the full description on the dataset page:
https://huggingface.co/datasets/tljohnsilver/zn-prompt-injection-bench.