Beta
Explore
Marketplace
Neural Labs
Playground
Wallet
Docs
d1_math_longest_3k_eval_636d – Dataset by mlfoundations-dev | AlphaNeural AI
You can deploy this model and start earning money today!
mlfoundations-dev
/
d1_math_longest_3k_eval_636d
like
0
1K<n<10K
parquet
tabular
text
datasets
dask
mlcroissant
polars
us
Views
No views yet
Model card
Files and Versions
Community
API
mlfoundations-dev/d1_math_longest_3k_eval_636d
Precomputed model outputs for evaluation.
Evaluation Results Summary
Metric AIME24 AMC23 MATH500 MMLUPro JEEBench GPQADiamond LiveCodeBench CodeElo CodeForces
Accuracy 25.7 66.5 79.4 30.4 47.0 45.5 28.8 7.8 8.2
AIME24
Average Accuracy: 25.67% ± 0.82% Number of Runs: 10
Run Accuracy Questions Solved Total Questions
1 30.00% 9 30
2 26.67% 8 30
3 26.67% 8 30
4 23.33% 7 30… See the full description on the dataset page:
https://huggingface.co/datasets/mlfoundations-dev/d1_math_longest_3k_eval_636d
.