A small, hand-labeled benchmark of code snippets — vulnerable, safe, and
needs-audit — for evaluating how well a tool detects security problems in
AI-generated code ("vibe coding"). Every case is a minimal, self-contained
example with a known, by-construction ground-truth label.
Crucially, the set is built around safe twins: many vulnerable cases are
paired with a near-identical safe variant living at the same file path. This
makes the benchmark… See the full description on the dataset page:
https://huggingface.co/datasets/axyr/ai-code-security-golden.