Views
No views yet
Qwen/Qwen3.8-2.4T-A95B with MoE layers quantized to NVFP4 and attention layers quantized to FP8 block1vllm serve RedHatAI/Qwen3.8-2.4T-A95B-NVFP4 \
2 --data-parallel-size 8 \
3 --enable-expert-parallel 8 \
4 --reasoning-parser qwen3 \
5 --max-num-seqs 140RedHatAI/Qwen3.8-2.4T-A95B-NVFP4.1inspect eval hf/Idavidrein/gpqa/diamond
2 --model vllm/RedHatAI/Qwen3.8-2.4T-A95B-NVFP4-FP8
3 --reasoning-effort xhigh
4 --model-base-url http://localhost:8000/v1
5 -M client_timeout=2400
6 --token-limit 100000
7 --retry-on-error=2| Benchmark | Qwen/Qwen3.8-2.4T-A95B | RedHatAI/Qwen3.8-2.4T-A95B-NVFP4-FP8 | RedHatAI/Qwen3.8-2.4T-A95B-NVFP4 | Inferact/Qwen3.8-2.4T-A95B-NVFP4 |
|---|---|---|---|---|
| GPQA DIamond | 92.6 | 93.1 | 92.9 | 92.9 |