Views
No views yet
d42ca8978c5a66e92c3446d46e8adfe03ef692ff. The same-source BF16 vision/video tower (333 tensors) and all 15 BF16 MTP tensors are retained.1hf download lyf/Qwen3.8-27B-Huihui-Abliterated-NVFP4-MTP-VL --local-dir ./huihui-qwen38-nvfp4
2
3docker run --rm --gpus all --ipc=host --network=host -e VLLM_NVFP4_GEMM_BACKEND=flashinfer-cutlass -e VLLM_USE_FLASHINFER_SAMPLER=1 -v "$PWD/huihui-qwen38-nvfp4:/model:ro" vllm/vllm-openai:qwen38-x86_64-cu130 /model --served-model-name qwen38-huihui-nvfp4 --host 0.0.0.0 --port 8000 --max-model-len 4096 --kv-cache-dtype fp8 --gpu-memory-utilization 0.92 --max-num-seqs 1 --max-num-batched-tokens 1024 --enable-prefix-caching --trust-remote-code --speculative-config '{"method":"qwen3_5_mtp","num_speculative_tokens":3}'| Component | Format/source |
|---|---|
| Language-model Linear layers | NVFP4 W4A4, group size 16 |
| Vision/video tower | 333 tensors, BF16, same Huihui checkpoint |
| MTP head | 15 tensors, BF16, same Huihui checkpoint |
lm_head, token embedding, GDN conv1d | BF16 |
| Calibration | CNN/DailyMail 3.0.0, 20 × 8192 tokens |
| Packaging | compressed-tensors nvfp4-pack-quantized |
vllm/vllm-openai:qwen38-x86_64-cu130:GET /health and /v1/models: passed0.602 / 0.429 / 0.295model-00001-of-00002.safetensors, model-00002-of-00002.safetensors: compressed checkpointmodel-mtp-extra.safetensors: 15 same-source BF16 MTP tensorsmodel.safetensors.index.json: complete 2687-tensor indexBUILD_MANIFEST.json, VALIDATION_REPORT.json, STATIC_VALIDATION_REPORT.json, recipe.yaml, SHA256SUMS: provenance and reproducibility