Per-round trajectory + per-category taxonomy from an automated red-teaming co-evolution loop (attacker -> target -> harm-judge) on Qwen2.5 (7B attacker+judge, 3B target), seeded with JailbreakBench behaviors. This dataset is the frozen-attacker control, contextual attack-specific refusals arm: held-in ASR flat ~0.45 across 5 rounds - even with proper contextual refusals, hardening does not reduce ASR.
Study question: does adversarial co-evolution ignite an arms… See the full description on the dataset page:
https://huggingface.co/datasets/yavuz-ai/seas-harden-ctx.