A code benchmark evaluation dataset with 83,072 solution trajectories generated by state-of-the-art thinking models on coding benchmark problems.
Each entry is a long-form solution trajectory (chain-of-thought + final code) produced by a reasoning model on a held-out coding benchmark. Every trajectory carries a verified correct label, and every problem carries a correct_ratio (pass rate over all trajectories for that problem).… See the full description on the dataset page:
https://huggingface.co/datasets/haowu89/math-ai-bench-sources-code.