Views
No views yet
bottlecapai/ThinkingCap-Qwen3.6-27B — BottleCap AI's token-efficient reasoning fine-tune of Qwen3.6-27B. Produced with llm-compressor → compressed-tensors, with the native MTP speculative-decode head preserved (bf16) and the Qwen3-VL vision tower preserved (bf16).<think> length by ~46 % on average (out-of-domain macro), and NVFP4 + the MTP draft cut the cost of every token. Fewer thinking tokens × faster tokens = a markedly snappier local reasoner — with no measurable accuracy loss on the parent's benchmarks (macro 81.5 → 80.7; LiveCodeBench 80.7 → 84.3; truncation 2.9 % → 0.4 %, per the base model card).--quantization flag needed (auto-detected).Qwen3_5ForConditionalGeneration (model_type qwen3_5), dense 27.8 B:full_attention_interval=4), hidden 5120, 262 K native context.--limit-mm-per-prompt.mtp_num_hidden_layers=1), kept bf16 → drives vLLM speculative decoding.<think>…</think>, use --reasoning-parser qwen3) — but a token-efficient one.QuantizationModifier(targets="Linear", scheme="NVFP4", # W4A4, group_size 16
ignore=["lm_head", "re:.*visual.*", "re:.*conv1d.*", "re:.*mtp.*"])conv1d, lm_head, and the entire MTP head stay bf16; everything else is NVFP4 W4A4. 32 calibration samples (neuralmagic/calibration), seq 8192, pure-CPU load (sequential-pipeline onload).model-base-aux.safetensors (15 bf16 tensors). Those are grafted into the NVFP4 output (model-mtp-bf16.safetensors) and spliced into the safetensors index.quantization_config.ignore, otherwise vLLM matches mtp.*_proj against targets=["Linear"], expects NVFP4 scales that do not exist, and loads the Qwen3_5MTP draft as garbage → 0 % spec-decode acceptance. This bake adds them automatically.1vllm serve sakamakismile/ThinkingCap-Qwen3.6-27B-NVFP4 \
2 --tensor-parallel-size 4 --max-model-len 131072 \
3 --max-num-seqs 16 --gpu-memory-utilization 0.90 --kv-cache-dtype fp8 \
4 --reasoning-parser qwen3 --limit-mm-per-prompt '{"image":0,"video":0}' \
5 --speculative-config '{"method":"qwen3_5_mtp","num_speculative_tokens":3}'NCCL_P2P_DISABLE=1 + --disable-custom-all-reduce (and NCCL_CUMEM_ENABLE=0 if TP=8 CUDA-graph capture hangs). Drop --speculative-config for plain decode. The hybrid model's KV is light (only the 16 full-attention layers cache), so full 128 K context fits even at TP=2.max_tokens ≥ 4096 (prefer 8192+). Even though ThinkingCap thinks less, at a tiny budget it can still spend it all inside <think> and return empty content.gptq_marlin_repack: size_n=24 not divisible by tile_n_size=64; the 24 attention-heads / DeltaNet odd dims violate Marlin's tile constraint). W4A4 avoids Marlin (NVFP4 cutlass/FlashInfer path).temperature=1.0, top_p=0.95, top_k=20.qwen3_5 dense+MTP recipe shared with sakamakismile/Qwen3.6-27B-MTP-pi-tune-NVFP4.