This latest (6-June-2026) NVFP4 quant of Qwen3.6-27B keeps quality very high, while improving
KLD, tail KLD metrics, top-token stability, and probability error compared to the earlier releases. It is faster and smaller in size.
It applies my RSF scale fitting technique to the Q_K quants and uses a different tensor mix layout compared to previously.
6-June-2026: changed the MTP tensors to be NVFP4, ~130tk/s tg is now possible in some configurations on 5090.
I am experimenting with a smaller verison of this model designed to be used on machines with 16GB of VRAM:
Qwen3.6-27B-NVFP4-SMALL-MTP-GGUF.
The SMALL version will be slower than this model but should maintain similar quality, although I am still evaluating it.
Feedback to improve this is appreciated.