This benchmark evaluation is for everyone. Open a discussion under "Community" to share your model's scores and average (with proofs) so we can officially put it on the leaderboard. Note that we only accept small models ranging from 0.5B to 3B.
Rank
Model
Params
🔢 gsm8krefn
💻 humanevalrefn
🧠 arcchalrefn
Overall Avg.