A benchmark dataset for evaluating automated code review agents against adversarial pull requests. The dataset contains 2,250 malicious PRs (current release: deterministic) and 347 benign ground-truth security fixes, grounded in real vulnerabilities from the OSV/SECommits database, across 10 CWE classes from the 2025 CWE Top 25.
Modern AI coding assistants can generate plausible-looking… See the full description on the dataset page:
https://huggingface.co/datasets/RedAI4Code/SEVRA.