Views
No views yet
bartowski/Qwen_Qwen3.6-35B-A3B-GGUF file
Qwen_Qwen3.6-35B-A3B-Q4_K_M.gguf.Qwen/Qwen3.6-35B-A3B repository was used for config,
tokenizer, chat template, and generation config only. Weight tensors come from
the Q4_K_M GGUF and were dequantized in fp32, transformed back to HF/vLLM
layout where llama.cpp stores Qwen3.5/3.6 Gated DeltaNet tensors differently,
then cast once to BF16.1VLLM_USE_FLASHINFER_SAMPLER=0 vllm serve . \
2 --served-model-name qwen36-gguf-q4km-bf16-gdnfix \
3 --dtype bfloat16 \
4 --max-model-len 4096 \
5 --moe-backend triton \
6 --attention-backend triton_attn \
7 --gdn-prefill-backend triton \
8 --trust-remote-code \
9 --language-model-onlyThe capital of France is begins with Paris.