AdvSafePrompt is a metadata-rich adversarial prompt dataset for LLM safety research.
The dataset extends the original Malicious LLM Prompts dataset using a rule-based adversarial augmentation framework.
It contains:
20,920 prompts
14 metadata fields
9 adversarial attack methods
Original + augmented prompts
Train / Validation / Test splits
Total Samples: 20,920
Original Samples: 5,098
Augmented Samples: 15… See the full description on the dataset page:
https://huggingface.co/datasets/tasnuvaesha2002/AdvSafePrompt.