ICE: Quantization by Error-Propagation Class in Sparse MoE Models
Technical report and full evidence for ICE (Isolation of Compounding
Error), a quantization bit-allocation method for sparse Mixture-of-Experts
checkpoints.
Every quantizer in the GGUF ecosystem minimizes the same objective for every
tensor: the importance-matrix-weighted error of that tensor's output, for the
current token. That is the right objective for a tensor whose error dies with
the token. It is the wrong objective for three other kinds, and no published
method separates them.
ICE classifies every tensor by how far its error travels:
class
mechanism
treatment
DISCRETE
error flips an argmax, so a different computation runs
exact (F32)
RECURRENT
error enters a state decay and compounds along the sequence
exact (F32)
CACHED
error is written to the KV cache once and re-read by every later token
near-exact (F16)
INSTANT
error affects this token only
this is where the budget lives
The first three are 0.14% of the model, so freezing them is a line item rather
than a trade-off. In one line: freeze what propagates, spend everything else on
the library.
Headline results
Mean KL divergence against the bf16 checkpoint, WikiText-2, one harness for all
files.
tier
size
mean KLD
nearest published tier
outcome
23G-ICE
22.83 GB
0.0361
UD-Q4_K_XL 23.21 GB / 0.0380
0.38 GB smaller, 5.0% better
23G-ICE
22.83 GB
0.0361
APEX-I-Quality 23.84 GB / 0.0415
1.01 GB smaller, 13.0% better
25G-ICE
24.84 GB
0.0303
APEX-I-Balanced 26.28 GB / 0.0345
1.44 GB smaller, 12.2% better
19G-ICE
18.82 GB
0.0608
UD-IQ4_XS 18.68 GB / 0.0723
15.8% better at +0.14 GB
Across the twelve-tier comparison, nine tiers are Pareto-optimal and three are
strictly dominated. ICE does not win at the top of the ladder: UD-Q5_K_S
and UD-Q6_K are the two best files measured, which the paper explains rather
than omits (Law 4).
The four laws
Sparsity. A bit on the always-on core is worth E/k bits on the expert
bank. Measured 31.8 against a predicted 32.
Convexity. Error falls as 4^-b, so allocation cleverness is capped at
+0.139 bpw. Two consequences: do no depth grading, and at a fixed average
always pick the narrower type bracket (measured +16.6% and +7.0% penalty
for widening).
Placement. Inside a fixed bracket, shallow-first is worth about −8.5%
KLD per bpw of gap and reverses below 0.47 bpw. Depth gain measured
directly as g(t) ≈ exp(−t/9.95).
The floor is epistemic. Fitted on two independent harnesses at k₀ =
0.0205 and 0.0219, and confirmed by a direct probe at 0.018297. The best file
measured is 8% above it. You cannot out-bit a wrong prior.
What is in this repository
paper/ICE_Technical_Report.md the report
appendix/original-ICE/ PRINCIPLE.md and RECIPE.md, the method as first written
appendix/recipes/ every tensor-type-file cited in the paper
appendix/measurements/ raw llama-perplexity output for every KLD quoted
appendix/SHA256SUMS.txt checksum for every file above
appendix/recipes/
suffix
meaning
_ICEbase
the current recipe, after the Law 3 revision
_previous
the recipe it replaced
_comparison
a published UD or APEX tier, as measured
cfg_21G-ICE_ICEbase.txt and cfg_21G-ICE_previous.txt have the same
SHA-256 (93456d45b1fc0dab...). That is not an oversight. 21G is the tier where
Law 3's condition does not hold, so the method's output is to change nothing, and
the identical checksum is the proof that nothing was changed.
Recipes for the three dominated tiers are not included: they were dropped after
the Pareto analysis and none was retained. Their measured numbers are in the
paper.
against a base.kld produced once from the bf16 checkpoint with the same corpus
and chunk count. Protocol details, including how the byte-identical control
variants are constructed, are in Appendix B.
The same recipe measured through three different paths gave 0.0345, 0.034162 and
0.034290, a spread under 1%.
Negative results
Seven are documented at the same weight as the positive ones, including the
retraction of a rule this work itself derived and shipped: bumping ffn_down
above its sibling projections, which is standard practice, measured 11.3% worse
than uniform at identical size. One of three registered predictions also failed,
and it is scored as such.
On the comparison ladders
Unsloth Dynamic 2.0 and LocalAI APEX are the work of their respective authors.
They are measured here, not reproduced or modified. The comparison exists because
no individual publication can provide it: each ladder is published with its own
harness and its own reference, so the tiers are not comparable until someone puts
them on one. The same analysis that finds ICE tiers dominating three others also
finds two UD tiers to be the best files on the board and two APEX tiers to be the
only options below 18.5 GB.
Citing
Cite this repository. Model cards for ICE-quantized GGUFs link here for the
method description.