Views
No views yet
moonshotai/Kimi-K3 with MoE layers quantized to NVFP4 for accelerated inference.1vllm serve RedHatAI/Kimi-K3-NVFP4 \
2 --tensor-parallel-size 8 \
3 --trust_remote_code \
4 --load-format instanttensor \
5 --reasoning-parser kimi_k3 \
6 --language-model-only # optional| Benchmark | moonshotai/Kimi-K3 | RedHatAI/Kimi-K3-NVFP4 |
|---|---|---|
| GPQA | 93.5 | 91.0 |
1inspect eval hf/Idavidrein/gpqa/diamond \
2 --model vllm/RedHatAI/Kimi-K3-NVFP4 \
3 --reasoning-effort high \
4 --model-base-url http://localhost:8000/v1