Views
No views yet
IQ4_KS (4.25 bpw quantization). The interesting part is that these models achieve a lower peplexity on
wiki.test.raw than the original bf16 model. This is surprising, considering that no QAT has been mentioned
in the Qwen3 announcement. Hence I'm putting them out there for anyone interested in evaluating performance by means other than PPL,
or just using for local inferrence.
For more details see this discussion.IQ4_KS quantization type is not available in mainline llama.cpp.