| Full-budget benchmark | Reference | Deimos R1 | Delta | Think tokens · ref → R1 |
|---|---|---|---|---|
| GSM8K · flexible | 0.860 | 0.907 | +0.047 | 1,778 → 357 |
| MMLU-Pro | 0.363 | 0.551 | +0.188 | 1,984 → 677 |
| IFEval · prompt loose | 0.260 | 0.353 | +0.093 | 4,808 → 1,189 |
| IFEval · instruction loose | 0.437 | 0.487 | +0.050 | — |
| IFEval · prompt strict | 0.260 | 0.267 | +0.007 | — |
| IFEval · instruction strict | 0.437 | 0.429 | −0.008 | — |
| GSM8K · strict format | 0.727 | 0.333 | −0.394 | — |
| 4,096-token constrained run | Reference | Deimos R1 | Delta |
|---|---|---|---|
| GSM8K · flexible | 0.660 | 0.933 | +0.273 |
| MMLU-Pro | 0.394 | 0.634 | +0.240 |
| IFEval · prompt strict | 0.247 | 0.320 | +0.073 |
#### N ending. Deimos R1 is materially weaker at silently copying that demonstrated format. Explicit format requests are more reliable.<think>…</think> before the final answer.vllm serve Michael-Kozu/Deimos-R1 --served-model-name deimos-r1 \
--max-model-len 8192 --gpu-memory-utilization 0.80 --trust-remote-code