This directory reproduces the Humanize + Comparator + AXLE experiment
The recorded run used gpt-5.5 with xhigh reasoning, a 50-turn cap, a
four-job probe, up to 64 concurrent workers, and a fallback concurrency of 16.
It produced 96 independently verified proofs. putnam_2013_a5 and
1970_a6 reached the turn cap and were not counted as proofs.
solve-all-putnambench.sh derives the… See the full description on the dataset page:
https://huggingface.co/datasets/humanfia-lab/putnambench-solution.