Views
No views yet
EigenLabs/Qwen3.8-27B-MTP-bf16
(@26a328e0) for the Layr-Labs qwen-3.8-mtp-challenge declared-head surface:
fc.weight at affine 6-bit group-64, the seven other 2D projections at
affine 4-bit group-64, 1D norms bf16. Produced with MLX 0.32.0 mx.quantize.
Per-tensor session measurements attribute the entire draft-acceptance cost of
uniform 4-bit quantization to fc.weight; 6-bit fc recovers roughly half of
it for +13.1 MB per draft step. The head only proposes — the pinned target
verifies every emitted token.