Experimental results from the CodeBench evaluation framework.
Two dataset configs are included.
Configs
h4_security — H4 Security-Adjusted Reliability Experiment
240 rollouts across 3 agents × 10 tasks × 8 rollouts.
Tasks split into 5 standard algorithmic tasks and 5 security-sensitive tasks
(eval, exec, shell, yaml, dynamic import).
agent_name
anote-code… See the full description on the dataset page:
https://huggingface.co/datasets/anote-ai/codebench-results.