A benchmark-aligned oQ4e (imatrix) → DWQ build of ThinkingCap-Qwen3.6-27B.
DWQ reduced held-out teacher divergence by 31% (0.0524 → 0.0362) while preserving
the original bf16 vision tower and native MTP speculative-decoding head.
A mixed-precision 4-bit MLX quantization of
bottlecapai/ThinkingCap-Qwen3.6-27B —
vision + text, with the original vision tower preserved in bf16 and the donor's
mixed-precision MTP head retained for multimodal inference and native speculative decoding in oMLX.
⚠️ These are quantized weights. The model's capability and efficient-thinking behaviour
come from BottleCapAI's base model — please star/cite it first. This repo contributes the
oQ/DWQ quantization recipe, benchmark-aligned calibration and MTP/vision packaging; DWQ tunes
the quantizer's scales/biases and does not add new knowledge.
Architecture: dense Qwen3.6-27B (qwen3_5 MLX architecture), 64 hybrid
Gated-DeltaNet/full-attention layers, up to 262,144-token context.
Vision tower: Qwen3.6 vision encoder, kept in bf16 and packaged as 333
vision_tower.* tensors.
ThinkingCap behaviour: BottleCapAI fine-tuned Qwen3.6-27B to preserve answer quality
while using roughly half as many thinking tokens on average. See the
base model card for its
multi-seed evaluation and full methodology.
Quantization: oQ4e (imatrix) → single-stage DWQ
This is not a plain round-to-nearest 4-bit conversion:
oQ4e — importance-aware mixed precision. oMLX's enhanced quantizer builds an
importance matrix from 1,024 × 512-token calibration samples and allocates additional
precision to sensitive tensors. The resulting config uses affine 4-bit, group size 64
as its base, with 110 quantized modules promoted to 5-bit and 2 modules to 6-bit.
DWQ — activation-aligned distillation. The trainable affine scales/biases of all
sub-8-bit language modules are optimized toward an oQ8e teacher made from the same base
model. The objective is KL divergence over the teacher's top-1024 logits at temperature 2.0.
Validation-first finalization. Training uses batch 1, 512-token windows, gradient
checkpointing, Adam, a cosine LR schedule (2.5e-7 peak, 50-step warmup, 0.1 end factor),
validation early stopping and exports only the best checkpoint.
Component splice. The donor's mixed-precision MTP head and bf16 vision tower are copied
back byte-for-byte after DWQ; neither is trained by the text-only DWQ pass.
Scheme
affine 4-bit base (g64) · imatrix-promoted 5/6-bit modules · one DWQ pass
The loss above measures fidelity to this recipe's 8-bit teacher on its held-out calibration
distribution. It is not task accuracy and should not be compared with losses from another
dataset, tokenizer, teacher or sequence length.
Calibration mix
The single DWQ stage retains general reasoning while emphasizing the model's intended code and
agent workloads. DWQ sees tokenized activation windows, not additional SFT updates.
Partition
Train / valid
Share
Sources / role
Resolved SWE agent trajectories
1,024 / 32
25%
SWE-smith tool calls, patches and terminal output
General reasoning
768 / 24
18.75%
OpenThoughts science/math/code traces
Verified math
512 / 16
12.5%
OpenR1-Math
Concise function code
512 / 16
12.5%
execution-filtered self-OSS-Instruct
Verified Python problems
512 / 16
12.5%
decontaminated/tested verifiable coding problems
Competitive code reasoning
512 / 16
12.5%
NVIDIA OpenCodeReasoning
General instruction/chat
256 / 8
6.25%
Tulu 3 SFT mixture
Agent + code data is 62.5% of the mix. Public multilingual SWE trajectories were not
available, so multilingual agent performance is inherited from the base rather than directly
represented by this calibration. The DWQ calibration itself is text-only; multimodal capability
comes from the unmodified bf16 vision tower and the distilled language backbone.
Evaluation
Quantization fidelity measured on this build
Metric
Raw oQ4e student
oQ4e → DWQ (this repo)
Held-out top-K teacher KL
0.0524
0.0362
Relative reduction
—
31%
No task or vision-benchmark score is claimed here yet. The result above demonstrates improved
language-logit fidelity on the held-out calibration split, not guaranteed benchmark improvement.
Inherited base-model results (bf16, not re-measured on this quant)
BottleCapAI reports the following for bottlecapai/ThinkingCap-Qwen3.6-27B with thinking
enabled and five seeds. These numbers describe the bf16 base, not this 4-bit build.
Benchmark
ThinkingCap bf16
Mean thinking-token reduction vs Qwen3.6-27B
GPQA-Diamond
83.8%
67.8%
MMLU-Pro
85.4%
53.7%
HMMT (Nov 2025)
84.7%
38.0%
LiveCodeBench
84.3%
41.1%
LongBench v2
60.2%
39.1%
RealWorldQA
81.9%
48.5%
GSM8K (in-domain holdout)
96.5%
74.1%
Harnesses and token budgets differ substantially, especially for LiveCodeBench. Consult the
source card
before comparing these figures with local evals.
The siblings share the same DWQ language backbone and MTP head; this repo additionally
contains the original bf16 vision tower.
How to run
These are MLX weights for Apple Silicon. The tested serving path is
oMLX, which supports Qwen3.6, vision input and native MTP speculative
decoding.
bash
1# Download directly into the oMLX model directory2hf download Litwein/ThinkingCap-Qwen3.6-27B-oQ4e-DWQ-MTP-Vision-MLX \3 --local-dir ~/.omlx/models/ThinkingCap-Qwen3.6-27B-oQ4e-DWQ-MTP-Vision-MLX
45# Enable native Lightning MTP once6curl -X PUT \7 http://127.0.0.1:8003/admin/api/models/ThinkingCap-Qwen3.6-27B-oQ4e-DWQ-MTP-Vision-MLX/settings \8 -H "Authorization: Bearer sk-local" -H "Content-Type: application/json"\9 -d '{"mtp_enabled": true}'1011# OpenAI-compatible multimodal API12curl -X POST http://127.0.0.1:8003/v1/chat/completions \13 -H "Authorization: Bearer sk-local" -H "Content-Type: application/json"\14 -d '{"model":"ThinkingCap-Qwen3.6-27B-oQ4e-DWQ-MTP-Vision-MLX",
15 "messages":[{"role":"user","content":[
16 {"type":"image_url","image_url":{"url":"https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}},
17 {"type":"text","text":"Describe the image and explain the visual clues."}
18 ]}],
19 "max_tokens":4096,"temperature":1.0,"top_p":0.95}'
The oMLX model id is case-sensitive and matches the downloaded folder name. Native MTP and
the multimodal packaging require a compatible runtime; oMLX is the path validated for this
release.
Recommended sampling
Use the base model's recommended settings: temperature=1.0, top_p=0.95, top_k=20, min_p=0.0 with thinking enabled. Give hard reasoning/code tasks a generous output budget
(32k or more where practical). Deterministic/greedy decoding is useful for reproducible
benchmarks, but is not the base author's recommended real-world sampling mode.
Intended use & limitations
Best suited to: multimodal reasoning, visual question answering, coding, tool/agent
workflows, math, STEM and long-context chat.
Quantization is lossy: for maximum fidelity use the bf16 base or a higher-bit quant.
DWQ is calibration, not SFT: it improves quantized-teacher fidelity on represented
activations; it does not teach facts or guarantee gains on every benchmark.
Vision was preserved, not vision-calibrated: the bf16 tower is unchanged, but this DWQ
run used text calibration and has not been assigned a measured vision-benchmark score.
Calibration coverage: public resolved SWE trajectories are Python-centric; multilingual
SWE/terminal/skills behaviour was not directly calibrated.
Long context costs memory: 262k is an architectural maximum, not a promise that every
Apple Silicon machine can allocate the corresponding KV/cache state.
Acknowledgements
BottleCapAI — creators of ThinkingCap and its
efficient-thinking finetune. All model capability comes from their base.
Qwen team — Qwen3.6-27B (Apache-2.0).
Apple MLX — mlx, mlx-lm, mlx-vlm and mlx_lm.quant.dwq.
Apache-2.0, inherited from bottlecapai/ThinkingCap-Qwen3.6-27B and Qwen3.6-27B.
Citation
Please cite the original ThinkingCap model:
bibtex
1@misc{ThinkingCap-Qwen3.6-27B,
2 title = {bottlecapai/ThinkingCap-Qwen3.6-27B},
3 author = {Lasocki, Karol and Osusky, Adam and Lindauer, Jan and Jirkovsky, Adam and
4 Mihal, Filip and Platek, Ondrej and Herel, David and Ihnatchenko, Luka and
5 Bartek, Vojtech and Jirak, Jiri and Mikolov, Tomas},
6 year = {2026}
7}