This repository contains imatrix-aware GGUF quantizations of the original Qwen3.6-27B MTP model using projected Mixture-of-Quantization (MoQ) tensor policies. These are quantizations of the base Qwen model, not a fine-tune or merge.
The complete local series was re-evaluated alongside kaitchup/Qwen3.6-27B-GGUF-MoQ under identical conditions. Lower is better in all three charts.
p999 KLD comparison
Mean KLD comparison
WikiText-2 PPL comparison
The interactive comparison report supports pan, wheel/mode-bar zoom, view reset, and cross-chart series toggles. The companion CSV contains all measured values and tensor-composition summaries.
Comparison Summary
All 18 measured GGUF files contain 866 tensors and 27,320,697,856 parameters. They were evaluated from scratch against the same BF16 reference logits on WikiText-2 with context length 512.
Across the 8 Jianqiao1 points inside the kaitchup size range, linearly interpolating the kaitchup curve at exactly the same file size favors Jianqiao1 on:
Mean KLD: 8 of 8 points
PPL: 8 of 8 points
p999 KLD: 8 of 8 points
4 especially close pairs favor Jianqiao1 on PPL, Mean KLD, and p999 KLD while the Jianqiao1 GGUF is slightly smaller:
Jianqiao1
GB
kaitchup
GB
PPL (J / K)
Mean KLD (J / K)
p999 KLD (J / K)
MoQ-3.8
12.836
MoQ-3.5
12.888
7.097372 / 7.410869
0.067854 / 0.103733
4.466393 / 4.815393
MoQ-4.1
14.249
MoQ-4.0
14.430
7.030544 / 7.077410
0.044025 / 0.054240
3.270178 / 3.549793
MoQ-4.6
15.242
MoQ-4.25
15.312
6.974555 / 7.030861
0.026953 / 0.038407
2.159290 / 2.421159
MoQ-4.8
16.150
MoQ-4.5
16.214
6.934147 / 6.992102
0.023306 / 0.027908
1.785062 / 2.034667
After refreshing the 3.2–4.1 recipes, the earlier local exceptions disappear: every Jianqiao1 point within the comparison range is now below the linearly interpolated kaitchup curve on all three headline metrics. The improvement is especially visible at 3.6 and 4.1, while 3.8 now forms a substantially stronger direct comparison with kaitchup 3.5.
Full Quality Results
Actual BPW is computed from the complete GGUF file size, including metadata and alignment. GB is decimal. PPL and all KLD values are lower-is-better; same top-p is higher-is-better.
Series
Recipe
Actual BPW
GB
PPL
Mean KLD
p999 KLD
p99 KLD
Max KLD
RMS delta-p
Same top-p
Jianqiao1
3.2
3.1649
10.808
7.446454
0.146377
6.369240
1.550732
22.504015
11.149%
84.226%
Jianqiao1
3.6
3.5547
12.140
7.249732
0.109670
5.453876
1.140977
23.917309
9.526%
86.368%
Jianqiao1
3.8
3.7586
12.836
7.097372
0.067854
4.466393
0.607818
26.374474
7.200%
89.505%
Jianqiao1
4.1
4.1725
14.249
7.030544
0.044025
3.270178
0.382479
24.835104
5.764%
91.354%
Jianqiao1
4.3
4.3929
15.002
7.019089
0.032922
2.589048
0.268183
20.440199
5.013%
92.337%
Jianqiao1
4.6
4.4631
15.242
6.974555
0.026953
2.159290
0.241018
21.444906
4.444%
93.530%
Jianqiao1
4.8
4.7291
16.150
6.934147
0.023306
1.785062
0.206606
19.002287
4.186%
94.007%
Jianqiao1
4.9
4.8385
16.524
6.951682
0.022715
1.951161
0.200312
22.226179
4.073%
94.172%
Jianqiao1
5.1
5.1102
17.452
6.921706
0.019116
1.664910
0.155483
23.198441
3.718%
94.671%
kaitchup
3.0
3.2891
11.232
7.731805
0.172659
6.413453
1.775455
24.286598
12.210%
82.250%
kaitchup
3.25
3.5421
12.097
7.553780
0.142644
5.771519
1.442072
23.057383
11.101%
83.508%
kaitchup
3.5
3.7739
12.888
7.410869
0.103733
4.815393
1.014503
24.127443
9.397%
85.579%
kaitchup
3.75
3.9913
13.631
7.150180
0.070493
4.349184
0.642432
23.034918
7.568%
89.194%
kaitchup
4.0
4.2255
14.430
7.077410
0.054240
3.549793
0.503634
23.122169
6.619%
90.449%
kaitchup
4.25
4.4838
15.312
7.030861
0.038407
2.421159
0.334282
24.549507
5.537%
91.695%
kaitchup
4.5
4.7478
16.214
6.992102
0.027908
2.034667
0.224925
22.857042
4.619%
92.787%
kaitchup
4.75
5.0489
17.243
7.032774
0.025216
1.862780
0.204153
22.090981
4.351%
93.166%
kaitchup
5.0
5.3188
18.164
7.033246
0.021426
1.817995
0.192830
20.425156
4.032%
94.184%
The kaitchup rows are comparison measurements only. kaitchup GGUF files are not redistributed in this repository.
Available Models
File
Actual BPW
Size GB
Size GiB
Qwen3.6-27B-MTP-MoQ-3.2.gguf
3.1649
10.808
10.066
Qwen3.6-27B-MTP-MoQ-3.6.gguf
3.5547
12.140
11.306
Qwen3.6-27B-MTP-MoQ-3.8.gguf
3.7586
12.836
11.955
Qwen3.6-27B-MTP-MoQ-4.1.gguf
4.1725
14.249
13.271
Qwen3.6-27B-MTP-MoQ-4.3.gguf
4.3929
15.002
13.972
Qwen3.6-27B-MTP-MoQ-4.6.gguf
4.4631
15.242
14.195
Qwen3.6-27B-MTP-MoQ-4.8.gguf
4.7291
16.150
15.041
Qwen3.6-27B-MTP-MoQ-4.9.gguf
4.8385
16.524
15.389
Qwen3.6-27B-MTP-MoQ-5.1.gguf
5.1102
17.452
16.253
Quantization Approach
The recipes in this repository are projected tensor policies derived from the observable assignments in w-ahmad/Qwen3.5-9B-GGUF-MoQ-MTP, adapted to Qwen3.6-27B MTP.
The process is:
Read tensor names, shapes, and GGML types from the source Qwen3.5-9B MoQ GGUF files.
Split tensor names into layer id and tensor suffix.
Map source and destination layers by normalized relative depth.
Reuse the source tensor type for matching suffixes, with a suffix-majority fallback where needed.
Apply an imatrix during quantization.
Keep normalization and other non-matmul tensors at their selected high precision.
Preserve MTP and force its eight large projection tensors in blk.64 to Q8_0.
This is a projection of observed MoQ policies, not the original unpublished optimizer.
All files in this repository were produced with the unsloth Qwen3.6-27B imatrix. A customized llama.cpp quantizer with per-tensor type assignments was used; no specific quantization command is required to use the resulting GGUF files.
Tensor Distribution
The counts below are retained from the existing master CSV. The refreshed 3.2–4.1 evaluator CSV does not include GGUF header metadata, so those tensor-composition counts were not recalculated in this update. Every model has 866 tensors in total.
Recipe
BF16
F32
IQ3_XXS
IQ3_S
IQ4_XS
Q2_K
Q3_K
Q4_K
Q5_K
Q8_0
3.2
0
360
256
0
17
224
1
0
0
8
3.6
0
360
128
240
18
112
0
0
0
8
3.8
0
360
64
256
82
96
0
0
0
8
4.1
0
360
64
64
258
96
0
16
0
8
4.3
0
360
16
0
258
96
0
128
0
8
4.6
0
360
0
0
352
0
0
128
18
8
4.8
48
360
0
0
288
0
0
80
82
8
4.9
0
360
0
0
256
0
0
112
130
8
5.1
96
360
0
0
176
0
0
32
194
8
Evaluation Conditions
The latest comparison used llama.cpp build 10122, commit d67c0b410, with CUDA on an RTX 5090. WikiText-2 wiki.test.raw was evaluated against BF16 logits saved from the original Qwen3.6-27B MTP GGUF. Context length remained fixed at 512; logical batch was 2048, ubatch was 8192, GPU layers were selected automatically with fit enabled, op offload and flash attention were enabled, and 16 CPU threads were used. All GPU evaluations were run strictly one at a time.
The BF16 reference PPL reported by the KLD evaluator was 6.902375.
Existing Performance Measurements
These throughput results were collected separately on an RTX 5090 and are retained as a practical reference. They are not part of the 2026-08-04 KLD comparison.
Model
pp512 tok/s
tg128 tok/s
pg32768,256 tok/s
MTP p512 prefill
MTP gen128
MTP p32768 prefill
MTP gen256
MoQ-4.8 MTP-Q8_0
2297.23
67.51
1793.82
1428.80
109.60
2270.90
90.10
MoQ-4.9 MTP-Q8_0
2273.29
66.83
1762.49
1459.60
102.30
2247.00
108.80
MoQ-5.1 MTP-Q8_0
2203.42
61.56
1716.57
1391.70
81.20
2209.70
87.50
Unsloth Q4_K_M
2217.93
65.52
1755.85
1265.20
94.80
2171.10
82.00
Usage
Use a recent llama.cpp build with Qwen3.6 MTP support. The files support ordinary generation and MTP speculative decoding. Choose a recipe according to available memory and the quality curves above; 4.8 is a strong balance point, while 5.1 provides the best measured overall quality in this set.
License and Acknowledgements
Released under the Apache License 2.0, following the base model license metadata.
Thanks to:
the Qwen team for Qwen3.6-27B and its MTP architecture;
w-ahmad for publishing the Qwen3.5-9B MoQ GGUF tensor policies used as the projection reference;
kaitchup for publishing an independent Qwen3.6-27B MoQ series that made this controlled comparison possible;
the unsloth team for the Qwen3.6-27B imatrix;
the llama.cpp project and contributors for GGUF quantization and evaluation tooling.