Views
No views yet
ModelOpt NVFP4_DEFAULT_CFGnvidia-modelopt via mtq.quantize and export_hf_checkpoint.512 samples from cnn_dailymail, sequence length 512,
batch size 2.1vllm serve rressl/GRM-2.6-Plus-NVFP4 \
2 --quantization modelopt \
3 --max-model-len 262144 \
4 --gpu-memory-utilization 0.90 \
5 --kv-cache-dtype fp8 \
6 --attention-backend flashinfer \
7 --enable-auto-tool-choice \
8 --tool-call-parser qwen3_coder \
9 --reasoning-parser qwen3 \
10 --trust-remote-code