Per-cell exploitation results from the V8 JavaScript engine benchmark, with full transcripts, tool-call logs, and capability grading. This dataset is the academic record for ExploitBench: succeeded runs and model-failed runs both ship, including cells where the model gamed the grader (see audit.json).
41 environments. Full list — one per env_id, sorted:
v8-crbug-1509576
v8-crbug-339064932… See the full description on the dataset page:
https://huggingface.co/datasets/exploitbench/v8.