Views
No views yet
| Benchmark | Datarus-R1-14B-Preview | QwQ-32B | Phi-4-reasoning | DeepSeek-R1-Distill-14B |
|---|---|---|---|---|
| LiveCodeBench v6 | 57.7 | 56.6 | 52.6 | 48.6 |
| AIME 2024 | 70.1 | 76.2 | 74.6* | - |
| AIME 2025 | 66.2 | 66.2 | 63.1* | - |
| GPQA Diamond | 62.1 | 60.1 | 55.0 | 58.6 |


<step>, <thought>, <action>, <action_input>, <observation> tags<think> and <answer> tags1@article{benchaliah2025datarus,
2 title={Datarus-R1: An Adaptive Multi-Step Reasoning LLM for Automated Data Analysis},
3 author={Ben Chaliah, Ayoub and Dellagi, Hela},
4 journal={arXiv preprint arXiv:2508.13382},
5 year={2025}
6}