Views
No views yet
Qwen3_5ForConditionalGeneration.pack-quantized format:
num_bits: 4, type: int, symmetric: true, group_size: 128,
strategy: groupignore),
so image understanding is preserved.1vllm serve <path-or-repo>/Qwen3.8-27B-Uncensored-W4A16 \
2 --host 0.0.0.0 --port 8000 \
3 --served-model-name qwen3.8-27b \
4 --gpu-memory-utilization 0.90 \
5 --max-model-len 32768 \
6 --enable-auto-tool-choice --tool-call-parser hermeshttp://localhost:8000/v1. Adjust
--max-model-len to trade context length against KV-cache memory on a 24 GB
card. A recent vLLM with compressed-tensors support is required.LICENSE. You must retain the license and attribution when redistributing.