Evaluation results for 8 models on AIME 2025 with 30 problems and 32 rollouts per problem.
We followed the evaluation guidelines and prompts from OLMo 3. Best effort was made to ensure reported numbers are as accurate as possible.
Code: pmahdavi/modal-eval
pmahdavi/Olmo-3.1-7B-Math-Code
73.3%
36.4%
2.4%… See the full description on the dataset page:
https://huggingface.co/datasets/pmahdavi/aime2025-merging-leaderboard.