Views
No views yet
dflash/ subfolder), transformed to run at tensor-parallel size 3 as the companion
drafter for
mitomtuna/MiMo-V2.5-0703-NVFP4-TP3
— the quantization of the exact target model this drafter was trained against (the
updated 2026-07-03 MiMo-V2.5 target that ships inside the DFlash repo).| axis | source | this repo |
|---|---|---|
| attention heads | 64 Q / 8 KV (head_dim 128) | 72 Q / 9 KV |
| MLP intermediate | 16384 | 16512 |
target-hidden fc, norms, mask embedding | unchanged | unchanged |
max_position_embeddings is set to 1,048,576 (extended from the shipped 262,144, matching
the common practice for this drafter; the rope bases are unchanged).--speculative-config model when serving the TP3 target with the
companion vLLM image — see the full command in the
target model card:{"model": "mitomtuna/MiMo-V2.5-DFlash-TP3", "method": "dflash", "num_speculative_tokens": 7}