This repo contains two experimental NVFP4 GGUF quantizations of
Qwen3.6-35B-A3B for
llama.cpp.
This was quantized using my experimental
advanced-gguf-quantizer tool.
Both models were imatrix calibrated for the first time using a new custom dataset that I am evaluating.
All PPL/KLD results were measured against the same BF16 wikitest KLD base, and then compared to the official NVFP4 release by NVIDIA.
Further evaluation tests are underway to identify real world performance differences between TURBO and HQ.