We present AutoBaxBuilder, an automated framework that generates code security benchmark tasks from scratch, reducing manual effort by ~12× while matching or outperforming expert tests and exploits.
Dataset Summary
AutoBaxBench is an agentically generated coding benchmark, generated by AutoBaxBuilder. It is designed to measure the ability of code generation models and agents to generate correct and secure code. The benchmark… See the full description on the dataset page: https://huggingface.co/datasets/eth-sri/AutoBaxBench.