The first labeled benchmark dataset for AI agent memory poisoning detection.
1,178 entries (856 clean, 322 poisoned) across 10 attack types, 7 domains, and 3 difficulty levels. Clean background text drawn from the AgentPoison (NeurIPS 2024) knowledge bases; poisoned entries authored for this dataset inspired by AgentPoison, MemoryGraft, and Microsoft advisory threat models.
There are 60+ prompt injection datasets. There are zero memory… See the full description on the dataset page:
https://huggingface.co/datasets/npow/memshield-bench.