Views
No views yet
| Setting | Value |
|---|---|
| Scheme | W4A16 (INT4 weights, FP16 activations) |
| Group size | 128 |
| Ignored layers | MoE gate layers (kept at full precision) |
| Method | RTN (iters=0) |
1vllm serve Lasimeri/MiniMax-M2.7-int4-AutoRound \
2 --trust-remote-code \
3 --tensor-parallel-size 8 \
4 --enable-auto-tool-choice \
5 --tool-call-parser minimax_m2 \
6 --reasoning-parser minimax_m2_append_think1python -m sglang.launch_server \
2 --model-path Lasimeri/MiniMax-M2.7-int4-AutoRound \
3 --trust-remote-code \
4 --tp 8 \
5 --reasoning-parser minimax-append-think \
6 --tool-call-parser minimax-m2| Component | Spec |
|---|---|
| CPU | AMD EPYC 7742 (64C / 128T) |
| RAM | 251 GB DDR4 |
| GPUs | 8× RTX 3080 (20 GB modded) |