Views
No views yet
| Property | Value |
|---|---|
| Base model | google/gemma-4-E4B-it |
| Quant method | NVIDIA ModelOpt (FP8 E4M3 - num_bits: (4, 3)) |
| Weight scheme | Per-channel (axis: 0) |
| Input activation | Dynamic Per-token (type: dynamic) |
| Calibration algorithm | max |
| Size | 12 GB (vs 15 GB BF16) |
modelopt quantization backend. Please ensure you refer to the vLLM documentation for Gemma 4 for advanced serving options.1vllm serve vrfai/gemma-4-E4B-it-fp8 \
2 --quantization modelopt