Views
No views yet
{1: 6630, 2: 2995, 3: 12263} (3-bit-heavy barbell).| corpus | KL(BF16 ‖ this build) | PPL BF16 → 2.4-bit |
|---|---|---|
| Japanese-reasoning hold-out | 0.219 | 11.51 → 12.14 |
| neutral multilingual (diverse) | 0.440 | 2.90 → 3.84 |
TRITON_ATTN
backend) + OneCompression VQ dequant kernels. No sm_120 sparse patch is needed — unlike DSA
models, M3's native indexer runs on stock kernels here. Tensor-parallel = 2, expert-parallel on.gpu_util 0.97 ≈ 372,736 tokens (~27 GiB/GPU free for KV —
the small weight footprint leaves generous headroom).<mm:think> blocks. The chat template exposes a thinking toggle —
chat_template_kwargs {"thinking_mode": "disabled"} gives direct answers; the default is
adaptive. Pair with vLLM's --reasoning-parser minimax_m3 to split reasoning from content.thinking_mode: "disabled" fully contains it.LICENSE). This is a quantized derivative and inherits that license.
Commercial use requires displaying "Built with MiniMax M3" and, above the revenue threshold
in the license, prior authorization from MiniMax — read LICENSE before any commercial
deployment.