Ablation-informed mixed-precision quantization of google/gemma-4-e2b-it. 3.0 GB file size, 139.69 perplexity — smaller than stock Q3_K_M (3.06 GB) with 24% lower perplexity.
Three ffn_gate layers identified by per-layer ablation as actively benefiting from Q2_K demotion. Not a blanket crush — surgical precision informed by 35-layer sensitivity sweep.
Benchmarks
Benchmark
Cerebellum v2
Q3_K_M Baseline
Delta
Perplexity (WikiText-2, 2048 ctx)
139.69
184.93
-24.4%
HumanEval pass@1
46.3%
46.3%
0.0%
ARC-Challenge
71.9%
71.9%
0.0%
HellaSwag
50.0%
50.0%
0.0%
MMLU-Redux
47.4%
47.6%
-0.2%
All benchmarks measured directly on this file. Identical benchmark performance at smaller size and significantly lower perplexity.
Why This Works
Standard quantization treats all layers identically. Cerebellum runs a per-layer ablation sweep — testing each layer's ffn_gate individually at Q2_K — and discovers that certain mid-network layers actually produce lower perplexity when crushed harder. This is a regularization effect: the gate tensors at layers 11, 13, and 14 (31-40% depth) carry redundant precision that creates noise at Q3_K_M.
The Regularization Effect
When we tested all 35 layers individually:
Layer
PPL at Q2_K
vs Baseline (184.93)
Effect
blk.11
169.18
-8.5%
Regularization
blk.13
170.71
-7.7%
Regularization
blk.14
172.70
-6.6%
Regularization
blk.12
176.27
-4.7%
Mild benefit
blk.0
169.21
-8.5%
Regularization
blk.30+
200+
+8%+
Damage
Layers 11, 13, 14 form a cluster in the mid-network where gate tensor precision actively hurts. Combining all three gives PPL 139.69 — the effects stack.
v1 vs v2: Proof That Precision Matters
Version
Method
PPL
HumanEval
ARC
HellaSwag
MMLU
Baseline
Stock Q3_K_M
184.93
46.3%
71.9%
50.0%
47.6%
v1
All 35 ffn_gate → Q2_K
139.34
17.7%
64.9%
39.9%
43.6%
v2
3 layers → Q2_K
139.69
46.3%
71.9%
50.0%
47.4%
v1 proved that blanket-crushing all ffn_gate tensors improves PPL but destroys benchmarks — same-layer interaction effects between simultaneously crushed tensors cause cascading damage. v2 proves that surgical, ablation-guided demotion captures the same PPL improvement with zero benchmark loss.
Architecture Family Recipe Transfer
This model validates that Cerebellum recipes transfer within architecture families. On Gemma 4 E4B (42 layers), the sweet spot for ffn_gate demotion is layers 14-17 (33-40% depth). On E2B (35 layers), it's layers 11-14 (31-40% depth). Same proportional position, same effect.
This means ablation results on one model in a family can inform the starting configuration for smaller/larger siblings — reducing the search space from O(layers) to O(1) confirmation.