Every run here is the same frozen contract: subset flash_10 (10 pairs / 10 tasks / 9 repos),
agent mini_swe_agent_v2, setting coop + --git, Modal sandboxes, step_limit=200,
max_model_len=49152. Produced by scripts/eval_e2e.py in
CooperTrain. Two runs are comparable only if their
run_meta.json agree — check it rather than assuming.
Each run directory contains, per pair:
agentN_traj.json
one agent's trajectory… See the full description on the dataset page:
https://huggingface.co/datasets/CooperBench/qwen35-9b-cooper-flash-evals.