Views
No views yet


| Model | Total Params | Active Params | Context Length | Precision | Download |
|---|---|---|---|---|---|
| MiMo-V2.5-Pro | 1.02T | 42B | 1M | FP8 (E4M3) Mixed | 🤗 HuggingFace 🤖 ModelScope |
| MiMo-V2.5-Pro-Base | 1.02T | 42B | 256K | FP8 (E4M3) Mixed | 🤗 HuggingFace 🤖 ModelScope |
| Category | Benchmark | Setting | MiMo-V2.5-Pro Base | MiMo-V2.5 Base | DeepSeek-V4-Pro Base | DeepSeek-V4-Flash Base | Kimi-K2 Base |
|---|---|---|---|---|---|---|---|
| Params | #Activated / #Total | - | 42B / 1.02T | 15B / 310B | 49B / 1.6T | 13B / 284B | 32B / 1.04T |
| General | BBH | 3-shot | 88.4 | 87.2 | 87.5 | 86.9 | 88.7 |
| MMLU | 5-shot | 89.4 | 86.3 | 90.1 | 88.7 | 87.8 | |
| MMLU-Redux | 5-shot | 92.8 | 89.8 | 90.8 | 89.4 | 90.2 | |
| MMLU-Pro | 5-shot | 68.5 | 65.8 | 73.5 | 68.3 | 69.2 | |
| DROP | 3-shot | 86.3 | 83.7 | 88.7 | 88.6 | 83.6 | |
| ARC-Challenge | 25-shot | 97.2 | 96.5 | - | - | 96.2 | |
| HellaSwag | 10-shot | 89.8 | 88.6 | 88.0 | 85.7 | 94.6 | |
| WinoGrande | 5-shot | 85.6 | 84.7 | 81.5 | 79.5 | 85.3 | |
| TriviaQA | 5-shot | 81.3 | 80.7 | 85.6 | 82.8 | 85.1 | |
| GPQA-Diamond | 5-shot | 66.7 | 58.1 | - | - | 48.1 | |
| Math | GSM8K | 8-shot | 99.6 | 83.3 | 92.6 | 90.8 | 92.1 |
| MATH | 4-shot | 86.2 | 67.7 | 64.5 | 57.4 | 70.2 | |
| AIME 24&25 | 2-shot | 37.3 | 36.9 | - | - | 31.6 | |
| Code | HumanEval+ | 1-shot | 75.6 | 71.3 | - | - | 84.8 |
| MBPP+ | 3-shot | 74.1 | 70.9 | - | - | 73.8 | |
| LiveCodeBench v6 | 1-shot | 39.6 | 35.5 | - | - | 26.3 | |
| SWE-Bench (AgentLess) | 3-shot | 35.7 | 30.8 | - | - | 28.2 | |
| Chinese | C-Eval | 5-shot | 91.5 | 88.6 | 93.1 | 92.1 | 92.5 |
| CMMLU | 5-shot | 90.2 | 88.2 | 90.8 | 90.4 | 90.9 | |
| Multilingual | GlobalMMLU | 5-shot | 83.6 | 77.4 | - | - | 80.7 |


| Component | MiMo-V2.5-Pro | MiMo-V2.5 |
|---|---|---|
| Total Parameters | 1.02T | 310B |
| Activated Parameters | 42B | 15B |
| Hidden Size | 6144 | 4096 |
| Num Layers | 70 (1 dense + 69 MoE) | 48 (1 dense + 47 MoE) |
| Full Attention Layers | 10 | 9 |
| SWA Layers | 60 | 39 |
| Num Attention Heads | 128 | 64 |
| Num KV Heads | 8 (GQA) | 8 (GA) / 4 (SWA) |
| Head Dim (QK / V) | 192 / 128 | 192 / 128 |
| Routed Experts | 384 | 256 |
| Experts per Token | 8 | 8 |
| MoE Intermediate Size | 2048 | 2048 |
| Dense Intermediate Size | 16384 (layer 0 only) | 16384 (layer 0 only) |
| SWA Window Size | 128 | 128 |
| Max Context Length | 1M | 1M |
| MTP Layers | 3 | 3 |
1SGLANG_ENABLE_SPEC_V2=1
2SGLANG_DEEPEP_NUM_MAX_DISPATCH_TOKENS_PER_RANK=256
3python3 -m sglang.launch_server \
4 --model-path XiaomiMiMo/MiMo-V2.5-Pro \
5 --trust-remote-code \
6 --pp-size 1 \
7 --dp-size 2 \
8 --ep-size 16 \
9 --tp-size 16 \
10 --moe-dense-tp-size 1 \
11 --enable-dp-attention \
12 --moe-a2a-backend deepep \
13 --dist-init-addr ${LWS_LEADER_IP}:20000 \
14 --node-rank ${LWS_WORKER_INDEX} \
15 --nnodes ${LWS_GROUP_SIZE} \
16 --page-size 64 \
17 --attention-backend fa3 \
18 --quantization fp8 \
19 --mem-fraction-static 0.7 \
20 --max-running-requests 128 \
21 --cuda-graph-max-bs 64 \
22 --chunked-prefill-size 32768 \
23 --context-length 1048576 \
24 --tokenizer-worker-num 64 \
25 --speculative-algorithm EAGLE \
26 --speculative-num-steps 3 \
27 --speculative-eagle-topk 1 \
28 --speculative-num-draft-tokens 4 \
29 --enable-multi-layer-eagle \
30 --host 0.0.0.0 \
31 --port 9001 \
32 --reasoning-parser mimo \
33 --tool-call-parser mimo \
34 --watchdog-timeout 3600 \
35 --model-loader-extra-config '{"enable_multithread_load": "true","num_threads": 64}'temperature=1.0, top_p=0.95.1@misc{mimo2026v25pro,
2 title={MiMo-V2.5-Pro},
3 author={{Xiaomi MiMo Team}},
4 year={2026},
5 howpublished={\url{https://huggingface.co/collections/XiaomiMiMo/mimo-v25}},
6}