Views
No views yet
tencent/Hy3 with MoE layers quantized to NVFP4 and attention layers quantized to FP8 block1vllm serve RedHatAI/Hy3-NVFP4-FP8 \
2 --tensor-parallel-size 4 \
3 --tool-call-parser hy_v3 \
4 --enable-auto-tool-choice \
5 --reasoning-parser hy_v3 \
6 --port 8089inspect eval hf/Idavidrein/gpqa/diamond --model vllm/RedHatAI/Hy3-NVFP4-FP8 --reasoning-effort high --model-base-url http://localhost:8089/v1