Views
No views yet

| Models | AIME24 avg@32 | AIME25 avg@32 | Minerva Math avg@4 | Olympiad Bench avg@4 | AMC23 avg@8 |
|---|---|---|---|---|---|
| DeepScaleR-1.5B | 43.1 | 27.2 | 34.6 | 40.7 | 50.6 |
| Qwen3-1.7B | 48.3 | 36.8 | 34.9 | 55.1 | 75.6 |
POLARIS-1.7B-Preview | 66.9 | 53.0 | 38.9 | 63.8 | 85.8 |
| Deepseek-R1-Distill-Qwen-7B | 55.0 | 39.7 | 36.7 | 56.8 | 81.9 |
| AReal-boba-RL-7B | 61.9 | 48.3 | 39.5 | 61.9 | 86.4 |
| Skywork-OR1-7B-Math | 69.8 | 52.3 | 40.8 | 63.2 | 85.3 |
POLARIS-7B-Preview | 72.6 | 52.6 | 40.2 | 65.4 | 89.0 |
| Deepseek-R1-Distill-Qwen-32B | 72.6 | 54.9 | 42.1 | 59.4 | 84.3 |
| qwen3-32B | 81.4 | 72.9 | 44.2 | 66.7 | 92.4 |
| qwen3-4B | 73.8 | 65.6 | 43.6 | 62.2 | 87.2 |
POLARIS-4B-Preview | 81.2 | 79.4 | 44.0 | 69.1 | 94.8 |
Qwen3-4B and DeepSeek-R1-Distill-Qwen-7B. Thanks for their wonderful work.1@misc{Polaris2025,
2 title = {POLARIS: A Post-Training Recipe for Scaling Reinforcement Learning on Advanced Reasoning Models},
3 url = {https://hkunlp.github.io/blog/2025/Polaris},
4 author = {An, Chenxin and Xie, Zhihui and Li, Xiaonan and Li, Lei and Zhang, Jun and Gong, Shansan and Zhong, Ming and Xu, Jingjing and Qiu, Xipeng and Wang, Mingxuan and Kong, Lingpeng}
5 year = {2025}
6}