Beta
Explore
Marketplace
Neural Labs
Playground
Wallet
Docs
d1_code_python_10k_eval_2e29 – Dataset by mlfoundations-dev | AlphaNeural AI
You can deploy this model and start earning money today!
mlfoundations-dev
/
d1_code_python_10k_eval_2e29
like
0
10K<n<100K
parquet
tabular
text
datasets
dask
mlcroissant
polars
us
Views
No views yet
Model card
Files and Versions
Community
API
mlfoundations-dev/d1_code_python_10k_eval_2e29
Precomputed model outputs for evaluation.
Evaluation Results Summary
Metric AIME24 AMC23 MATH500 MMLUPro JEEBench GPQADiamond LiveCodeBench CodeElo CodeForces AIME25 HLE LiveCodeBenchv5
Accuracy 16.7 62.7 75.8 29.2 46.3 38.6 40.4 11.9 15.2 13.0 9.1 28.2
AIME24
Average Accuracy: 16.67% ± 1.33% Number of Runs: 10
Run Accuracy Questions Solved Total Questions
1 20.00% 6 30… See the full description on the dataset page:
https://huggingface.co/datasets/mlfoundations-dev/d1_code_python_10k_eval_2e29
.