Views
No views yet
google/gemma-4-12B-it-assistant,
the Multi-Token Prediction (MTP) drafter that pairs with the Gemma 4 12B-it
target for speculative decoding. The drafter is a small 4-layer model; its
linear layers are quantized to FP8 (E4M3) with per-tensor static scales via
NVIDIA ModelOpt. The
drafter↔target handshake projections (pre_projection, post_projection) and
lm_head stay in BF16.shared_kv_states at inference time. Use it as a spec model in
vLLM, paired with the 12B-it target.bahadirakdemir/gemma-4-12B-it-text-fp8 —
the FP8-quantized 12B-it text tower produced by the same pipeline. The FP8
scales were calibrated by running real speculative decoding against the 12B-it
target over 32 instruct-style prompts.gemma4_unified), which is newer than the classic gemma4 (e.g. 31B). You need:gemma4_unified support — at the time of writing this is on the
main branch / nightly (uv pip install -U vllm --pre), not yet in a tagged
stable release (≤ 0.22.0). It will be in the next stable release.1vllm serve bahadirakdemir/gemma-4-12B-it-text-fp8 \
2 --quantization modelopt \
3 --max-model-len 8192 \
4 --max-num-batched-tokens 8192 \
5 --gpu-memory-utilization 0.5 \
6 --limit-mm-per-prompt '{"image": 0, "audio": 0}' \
7 --speculative-config '{"model": "bahadirakdemir/gemma-4-12B-it-assistant-fp8", "num_speculative_tokens": 4}'vllm/vllm-openai:gemma4-0505-arm64-cu130 on NVIDIA GB10.bahadirakdemir/gemma-4-31B-it-assistant-fp8,
produced by the same pipeline.