Laguna-S-2.1 — imatrix calibration study
Importance matrices (imatrices) for
poolside/Laguna-S-2.1 (118B-total MoE) computed from
five different calibration corpora, published together with the corpora themselves, to map how calibration
domain affects the imatrix — and, downstream, low-bit quant quality on a coding model.
Most published imatrices use one general-purpose corpus with some code mixed in (e.g. bartowski's, unsloth's). A deliberately code-heavy or single-language calibration for a coding model is largely unexplored, with no public head-to-head. This repo is that experiment.
Why calibration domain matters here
An imatrix records per-input-channel activation energy (Σx²) so the quantizer spends its bit budget where it reduces output error most. It's data-dependent: the corpus is the domain-specialization knob. On this MoE specifically, we measured that the routed per-expert tensors (ffn_*_exps) are by far the most calibration-sensitive, while attention / router / shared-expert / dense tensors are corpus-robust. Those routed experts are exactly what gets pushed to 2–3 bits in the aggressive APEX tiers — so on a coding model, a code-matched imatrix should show up most in the small quants.
Imatrices
All five were produced by the same forward-validated band-serialized torch generator (bf16, ctx 512, 126 chunks), which matches a reference transformers LagunaModel forward to bf16 precision (per-layer corr > 0.9999). Each covers all 572 ggml tensors.
| imatrix | calibration corpus | domain | corpus license |
|---|
imatrices/laguna-datav3.imatrix | corpora/laguna_calib_datav3.txt | general/scientific prose, no code (baseline) | permissive |
imatrices/laguna-codeweighted.imatrix | corpora/laguna_calib_codeweighted.txt | ~70% multi-language code + 30% prose | permissive |
imatrices/laguna-purecode.imatrix | corpora/laguna_calib_purecode.txt | 100% code, 8 languages (Python, Go, C, C++, CUDA, JS, TS, Rust) | permissive |
imatrices/laguna-python.imatrix | corpora/laguna_calib_python.txt | 100% Python | permissive |
imatrices/laguna-golang.imatrix | corpora/laguna_calib_golang.txt | 100% Go | permissive |
imatrices/laguna-webdev.imatrix | corpora/laguna_calib_webdev.txt | web front-end: HTML + CSS/SCSS + JS + TS | permissive |
imatrices/laguna-rust.imatrix | corpora/laguna_calib_rust.txt | 100% Rust | permissive |
imatrices/laguna-kernel.imatrix | corpora/laguna_calib_kernel.txt | 100% Linux-kernel C | GPL-2.0 ⚠️ |
imatrices/laguna-literature.imatrix | corpora/laguna_calib_literature.txt | classic English literature (no code) | public domain |
⚠️ Mixed licensing. Every corpus except kernel is assembled from permissively-licensed source (MIT / Apache-2.0 / BSD / Unlicense / public-domain prose). The kernel corpus is verbatim Linux-kernel C and is therefore GPL-2.0 — it is redistributed here under GPL-2.0 with attribution. If you need a clean permissive-only subset, take everything except laguna_calib_kernel.txt. The imatrix files are aggregate statistics (Σx² per channel), not derivative works of any source code, and are provided under this repo's MIT license regardless of which corpus produced them.
Findings
Median per-tensor correlation of each imatrix vs the datav3 no-code baseline, by tensor category. router sees every token, so it's the corpus-robustness reference; *_exps are the routed experts (the tensors crushed to 2–3 bits in the aggressive quant tiers). Lower = the calibration moved that group more.
| corpus | attn | router | gateup_exps | down_exps | overall |
|---|
| literature (prose control) | 0.981 | 0.936 | 0.696 | 0.562 | 0.959 |
| golang | 0.973 | 0.903 | 0.702 | 0.583 | 0.937 |
| purecode (8-lang) | 0.971 | 0.898 | 0.689 | 0.600 | 0.933 |
| python | 0.973 | 0.891 | 0.683 | 0.519 | 0.931 |
| rust | 0.969 | 0.882 | 0.681 | 0.568 | 0.931 |
| kernel (C) | 0.972 | 0.889 | 0.668 | 0.480 | 0.929 |
| webdev | 0.970 | 0.891 | 0.659 | 0.541 | 0.927 |
| codeweighted | 0.965 | 0.882 | 0.673 | 0.557 | 0.912 |
1. Code vs prose is the dominant axis — and it's a real signal, not noise. literature is a completely different prose domain from datav3 (Victorian novels vs scientific abstracts), yet it sits clearly above every code corpus: router 0.936 / overall 0.959, vs 0.88–0.90 / 0.91–0.94 for anything containing code. A pure-noise explanation would put literature down with the code corpora; instead it stays with prose. So calibrating on code genuinely shifts the imatrix in a code-specific direction.
2. Single language is a measurable secondary axis. Direct language-vs-language, expert tensors only:
| pair | gateup_exps | down_exps | overall |
|---|
| python ↔ golang | 0.721 | 0.608 | 0.980 |
| python ↔ rust | 0.695 | 0.602 | 0.977 |
| python ↔ kernel-C | 0.696 | 0.644 | 0.971 |
Different languages specialize the expert importance measurably (experts drop to ~0.60–0.72) while staying ~0.97–0.98 overall. Kernel-C is the most Python-divergent — systems C vs high-level Python.
3. The shift lives in the routed experts. Across every corpus, attn (0.96–0.98) and router (0.88–0.94) stay robust while down_exps falls to 0.48–0.60 and gateup_exps to ~0.66–0.70. Those routed experts are exactly the tensors quantized to IQ2_S / IQ3_XXS / Q3_K in the low-bit APEX tiers — so on a coding model, a code-matched imatrix should help most precisely where the small quants are most fragile.
Methodology: torch band-serialized generator, bf16, ctx 512 × 126 chunks per corpus; correlation computed on per-column mean activation energy (Σx²/count).
Using an imatrix
1llama-quantize --tensor-type-file <apex-config>.txt \
2 --imatrix laguna-purecode.imatrix \
3 Laguna-S-2.1-BF16-00001-of-00005.gguf \
4 Laguna-S-2.1-APEX-i-mini-purecode.gguf Q6_K
Requires a llama.cpp build with the laguna architecture (e.g. the poolside fork). The imatrix helps most for the low-bit tiers (i-mini / i-subcompact), where the routed experts are quantized aggressively.
Corpora provenance & licenses
The calibration corpora are small (~0.5 MB each) concatenations of source text sampled for calibration, from public repositories only — no private/proprietary code. Redistributed with attribution under each source's original license:
- Python: HuggingFace
transformers (Apache-2.0), llama.cpp (MIT), apex-quant (MIT).
- C/C++/CUDA:
llama.cpp / poolside fork (MIT).
- Go: gin-gonic/gin (MIT), spf13/cobra (Apache-2.0), prometheus/client_golang (Apache-2.0).
- Rust: BurntSushi/ripgrep (Unlicense), serde-rs/serde (MIT/Apache-2.0).
- JS/TS: expressjs/express (MIT), colinhacks/zod (MIT).
- Web (HTML/CSS/SCSS): twbs/bootstrap (MIT), h5bp/html5-boilerplate (MIT).
- Linux kernel C: torvalds/linux — GPL-2.0 (mm/, kernel/, lib/, fs/, net/ subsystems). This is the only non-permissive corpus.
- Literature: classic out-of-copyright English novels via Project Gutenberg (Austen, Melville, Doyle, Shelley, Dickens, Carroll, Twain) — public domain; Gutenberg header/footer boilerplate stripped.
- Prose:
datav3 / codeweighted prose is general scientific abstracts + wikitext-style text.
No test files were included; minified/generated blobs were filtered out. This repo does not relicense the underlying source — each sample remains under its original license; see the upstream projects for full terms.
Reproduce
Generator: band-serialized torch imatrix tool for the laguna arch (serialized_gen_laguna.py), run per corpus:
1python serialized_gen_laguna.py --model <laguna-hf-dir> \
2 --calib corpora/laguna_calib_python.txt \
3 --out laguna-python.imatrix \
4 --ctx 512 --chunks 126 --band 4 --batch 8 --device cuda