Views
No views yet
transformers==5.3.0llm-compressor==0.14.1.dev24vllm==0.17.1| Precision | Layers |
|---|---|
| FP8 W8A8 | most Linear layers (per-channel static weight scales, per-token dynamic input scales) |
| BF16 | lm_head, embed_tokens, self_attn.o_proj, DeltaNet linear_attn.out_proj, DeltaNet in_proj_a/in_proj_b, visual encoder, MTP sidecar |
model_type=qwen3_564 text layers (hybrid DeltaNet + softmax, full_attention_interval=4)mtp_num_hidden_layers=1max_position_embeddings=262144hidden_size=5120, intermediate_size=17408vocab_size=248320pip install -U vllm>=0.17.0 transformers>=5.3.01vllm serve mconcat/Qwopus3.5-27B-v3-FP8-Dynamic \
2 --max-model-len 32768 \
3 --gpu-memory-utilization 0.85 \
4 --max-num-seqs 1 \
5 --skip-mm-profiling \
6 --reasoning-parser qwen31vllm serve mconcat/Qwopus3.5-27B-v3-FP8-Dynamic \
2 --max-model-len 32768 \
3 --gpu-memory-utilization 0.85 \
4 --max-num-seqs 1 \
5 --skip-mm-profiling \
6 --reasoning-parser qwen3 \
7 --speculative-config '{"method":"mtp","num_speculative_tokens":1}'1from transformers import AutoTokenizer, Qwen3_5ForConditionalGeneration
2import torch
3
4model = Qwen3_5ForConditionalGeneration.from_pretrained(
5 "mconcat/Qwopus3.5-27B-v3-FP8-Dynamic",
6 torch_dtype=torch.bfloat16,
7 device_map="auto",
8 trust_remote_code=True,
9)
10
11tokenizer = AutoTokenizer.from_pretrained(
12 "mconcat/Qwopus3.5-27B-v3-FP8-Dynamic",
13 trust_remote_code=True,
14)| Framework | Supported | Notes |
|---|---|---|
| vLLM >= 0.17.0 | Yes | Verified with vllm==0.17.1 on Blackwell; MTP works |
| transformers >= 5.3.0 | Yes | Direct loading with device_map="auto" |
| SGLang | Unknown | Not verified |
self_attn.o_proj and DeltaNet linear_attn.out_proj in BF16 to preserve output projection fidelity.model.safetensors file (no separate model.mtp.safetensors).--skip-mm-profiling with vLLM to skip vision encoder profiling.>= 9 to 9 <= x < 12 in vllm/model_executor/layers/fla/ops/utils.py.--kv-cache-dtype fp8_e4m3 with this model family — the checkpoint lacks calibrated KV scales and will produce degraded output. Use the default BF16 KV cache.