mlfoundations-dev/openthoughts3_100k_eval_08c7
Precomputed model outputs for evaluation.
Average Accuracy: 29.33% ± 2.25%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
23.33%
7
30
2
16.67%
5
30
3
30.00%
9
30
4
36.67%
11
30
5
36.67%
11
30
6
33.33%
10
30
7
40.00%
12
30
8
30.00%
9
30
9
23.33%
7
30
10
23.33%
7
30