Views
No views yet
Enterprise-grade OCP FP8 quantized Qwen-3 14B for AMD ROCm, end-to-end KV-cache in FP8 with Quark
Linear layers (excluding lm_head), activations, and KV cache| Metric | FP16 Baseline | FP8_e4m3 Quantized |
|---|---|---|
| Wikitext2 Perplexity | 6.38 | 8.73 |
| Memory Footprint | 1.0× | 0.56× |