Views
No views yet
fp8 quantization of Qwen/Qwen3.8-2.4T-A95B, produced with compressed-tensors by streaming the checkpoint tensor-by-tensor (the model is never fully instantiated).float8_e4m3fn), per output channel, with dynamic per-token activation quantization. Highest fidelity of the set and the largest.| Variant | Format | Size | vs BF16 | Mean rel. error | Linears quantized | Left BF16 |
|---|---|---|---|---|---|---|
| Qwen3.8-2.4T-A95B-FP8 ← this one | float-quantized | 2453.05 GB | 50% | 0.0264 | 143569 | 0 |
| Qwen3.8-2.4T-A95B-NVFP4 | nvfp4-pack-quantized | 1382.45 GB | 28% | 0.0952 | 143569 | 0 |
| Qwen3.8-2.4T-A95B-int4 | pack-quantized | 1268.00 GB | 26% | 0.1118 | 143569 | 0 |
||dequant(W) - W|| / ||W||, averaged over a sample of quantized Linear layers, measured against the original BF16 weights. Lower is better.| Format | float-quantized |
| Weight bits | 8 |
| Group size | per-channel (no grouping) |
| Strategy | channel |
| Linears quantized | 143569 |
| Left in BF16 | 0 |
| Shards | 572 |
| On disk | 2453.05 GB |
| Mean relative error | 0.0264 |
| Shape/dtype conformance failures | 0 |
vllm serve dudeman2512/Qwen3.8-2.4T-A95B-FP8