Views
No views yet
14e9d62, Foundry 2f99202. Search ledger and the full
pre-registered comparison record live with the run.| file | size | PPL (wikitext-2, ctx 512, full corpus) | vs BF16 (6.83) |
|---|---|---|---|
...-Q4_K_M.gguf | 17.60 GiB | 6.8948 ± 0.046 | +0.95% |
...-Q5_K_M.gguf | 20.60 GiB | 6.8579 | +0.41% |
...-Q6_K.gguf | 24.13 GiB | 6.7979 | −0.47% (within noise of baseline) |
D:Q4_K_M E:Q8_0 H:BF16 K:BF16 O:BF16 Q:Q8_0 S:Q8_0 U:IQ4_NL X:MXFP4_MOE).At matched 17.6 GiB, the measured per-group search beat the predicted per-tensor allocation: +0.95% vs +3.47% quality loss — a +2.50% PPL gap, far outside the measurement error.
never-quantize and f32-required-operand classes); BF16-designated groups
are written as F16 on disk (llama.cpp BF16 compute-graph limitation).