Views
No views yet
RedHatAI/Qwen3.6-35B-A3B-NVFP4: NVFP4 4-bit weights + activations on the
language model; vision tower, lm_head, and MTP kept in high precision for accuracy.compressed-tensors / nvfp4-pack-quantized → served directly by vLLM.1vllm serve mkd-hossain/Keural-Nova-v1.2-experimental-NVFP4 \
2 --served-model-name Keural-Nova-v1.2 \
3 --tensor-parallel-size 2 \
4 --max-model-len 262144 \
5 --tool-call-parser qwen3_xml --enable-auto-tool-choicellm-compressor (compressed-tensors), NVFP4 scheme.lm_head + MTP excluded (bf16).nvidia/RedHatAI NVFP4 Qwen3.6-A3B checkpoints. Validate on your Blackwell against the
bf16 base (Keural-Nova-v1.2-experimental)
on your workload (tool-calling, Korean, code) before production.RedHatAI/Qwen3.6-35B-A3B-NVFP4. "Keural Nova" is a model by MKD.