Views
No views yet
self_attn, and the GatedDeltaNet out_proj.in_proj_a/in_proj_b, lm_head, embeddings, vision tower, MTP head.compressed-tensors (pack-quantized). Full recipe: recipe.yaml.| Benchmark | Score |
|---|---|
| GSM8K | 96.8% (242/250) |
| MMLU-Pro | 82.4% (412/500) |
1# W4A16 — int4 weights, fp16 activations
2vllm serve Avesed/Qwen3.6-27B-INT4-W4A16 \
3 --tensor-parallel-size 2 --trust-remote-code --reasoning-parser qwen31vllm serve Avesed/Qwen3.6-27B-INT4-W4A16 \
2 --tensor-parallel-size 2 --marlin-input-dtype int8 --trust-remote-code --reasoning-parser qwen3