mlfoundations-dev/openmathreasoning_30k_eval_08c7
Precomputed model outputs for evaluation.
Average Accuracy: 15.33% ± 1.07%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
16.67%
5
30
2
13.33%
4
30
3
20.00%
6
30
4
13.33%
4
30
5
10.00%
3
30
6
20.00%
6
30
7
13.33%
4
30
8
13.33%
4
30
9
13.33%
4
30
10
20.00%
6
30