mlfoundations-dev/openthoughts3_math_100k_eval_08c7
Precomputed model outputs for evaluation.
Average Accuracy: 13.33% ± 1.41%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
13.33%
4
30
2
13.33%
4
30
3
16.67%
5
30
4
10.00%
3
30
5
20.00%
6
30
6
16.67%
5
30
7
13.33%
4
30
8
16.67%
5
30
9
10.00%
3
30
10
3.33%
1
30