Views
No views yet
embed_tokens and lm_head (vocab 200,064) — not shared with
the target. Everything is RTN-quantized to int4 (compressed-tensors
pack-quantized, group 128): Linear, embedding, and lm_head. ~1.6 GB.--speculative-config '{"method":"eagle3","model":"Sebesky/MiniMax-M3-EAGLE3-RTN-INT4","num_speculative_tokens":3}'draft_tensor_parallel_size: 2:
acceptance length 3.36 (78.6% draft acceptance) at 120k-token context,
decode 39.7 / 30.4 / 26.4 tok/s @512/65k/120k with 4-bit KV quantization.