Views
No views yet
b47b90d0be80981f6f476c997afbdfab99bba6c7.
At the time of quantization, the support branch commit was
be7d53ceec0c823c07d540e21828d7d187b02d3b.5bb2892a899809fd903829735f927207da358f2f.| Quant | Status | Size | PPL | KL divergence | Same top |
|---|---|---|---|---|---|
IQ1_S_R4 | uploaded | 43.54 GiB | 17.029677 +/- 0.150383 | 1.771976 +/- 0.006119 | 55.882% +/- 0.137 |
IQ1_S | uploaded | 43.91 GiB | 20.451320 +/- 0.194807 | 1.920889 +/- 0.006747 | 53.988% +/- 0.138 |
IQ2_XS | uploaded | 63.07 GiB | 12.307893 +/- 0.111532 | 1.325747 +/- 0.005793 | 63.043% +/- 0.133 |
IQ2_S | uploaded | 63.08 GiB | 12.470802 +/- 0.113326 | 1.340981 +/- 0.005827 | 62.743% +/- 0.134 |
IQ2_K_R4 | uploaded | 64.35 GiB | 13.397426 +/- 0.128657 | 1.387314 +/- 0.006164 | 62.577% +/- 0.134 |
IQ2_M | uploaded | 69.47 GiB | 12.101664 +/- 0.109158 | 1.294993 +/- 0.005730 | 63.149% +/- 0.133 |
Q2_K | uploaded | 77.66 GiB | 12.881872 +/- 0.123074 | 1.330646 +/- 0.005951 | 63.376% +/- 0.133 |
IQ3_XXS | uploaded | 81.98 GiB | 11.366560 +/- 0.103969 | 1.217616 +/- 0.005754 | 65.037% +/- 0.132 |
IQ3_S | uploaded | 91.81 GiB | 10.821543 +/- 0.095954 | 1.169696 +/- 0.005503 | 65.583% +/- 0.131 |
IQ3_K_R4 | uploaded | 91.81 GiB | 11.210677 +/- 0.102925 | 1.181941 +/- 0.005637 | 65.717% +/- 0.131 |
Q3_K_S | uploaded | 91.81 GiB | 11.531644 +/- 0.106822 | 1.214226 +/- 0.005712 | 65.297% +/- 0.132 |
Q3_K_M | uploaded | 100.91 GiB | 11.246193 +/- 0.103685 | 1.186657 +/- 0.005657 | 65.816% +/- 0.131 |
IQ4_XS | uploaded | 112.80 GiB | 4.615172 +/- 0.033625 | 0.067625 +/- 0.000706 | 88.899% +/- 0.087 |
IQ4_KS_R4 | uploaded | 113.20 GiB | 4.598246 +/- 0.033452 | 0.063984 +/- 0.000683 | 89.217% +/- 0.086 |
IQ4_NL | uploaded | 119.26 GiB | 4.605535 +/- 0.033524 | 0.066829 +/- 0.000720 | 89.049% +/- 0.086 |
IQ4_K_R4 | uploaded | 119.26 GiB | 4.617317 +/- 0.033749 | 0.058367 +/- 0.000667 | 89.440% +/- 0.085 |
Q4_K_S | uploaded | 119.91 GiB | 4.599690 +/- 0.033491 | 0.065115 +/- 0.000694 | 89.013% +/- 0.086 |
Q4_K_M | uploaded | 127.33 GiB | 4.602518 +/- 0.033538 | 0.061163 +/- 0.000667 | 89.209% +/- 0.086 |
IQ5_KS_R4 | uploaded | 139.04 GiB | 4.608804 +/- 0.033738 | 0.046126 +/- 0.000597 | 90.056% +/- 0.083 |
IQ5_KS | uploaded | 139.04 GiB | 4.608804 +/- 0.033738 | 0.046126 +/- 0.000597 | 90.056% +/- 0.083 |
Q5_K_S | uploaded | 145.10 GiB | 4.598543 +/- 0.033604 | 0.048429 +/- 0.000630 | 89.865% +/- 0.083 |
IQ5_K_R4 | uploaded | 145.10 GiB | 4.624008 +/- 0.033904 | 0.041238 +/- 0.000541 | 90.446% +/- 0.081 |
IQ5_K | uploaded | 145.10 GiB | 4.624008 +/- 0.033904 | 0.041238 +/- 0.000541 | 90.446% +/- 0.081 |
Q5_K_M | uploaded | 149.26 GiB | 4.607405 +/- 0.033688 | 0.045827 +/- 0.000562 | 90.036% +/- 0.083 |
IQ6_K | uploaded | 174.48 GiB | 4.633778 +/- 0.033986 | 0.028084 +/- 0.000440 | 91.120% +/- 0.079 |
Q6_K | uploaded | 172.85 GiB | 4.647079 +/- 0.034139 | 0.029602 +/- 0.000429 | 90.977% +/- 0.079 |
Q8_0 | uploaded | 223.63 GiB | 4.633481 +/- 0.033995 | 0.022906 +/- 0.000376 | 91.302% +/- 0.078 |
BF16 | uploaded | 420.71 GiB | 4.683023 +/- 0.034506 | baseline | baseline |
laguna-m1-q8-public-v1-256-gatefix.imatrix449427a1a4c5d2480d98e3fcccc3884bd3aafa325d937741a162adf84e7e7eed2.6084 +/- 0.00786llama-imatrix from
Laguna-M.1-Q8_0-00001-of-00010.gguf using -c 4096, -ngl 999,
--chunks 256, --no-warmup, and all available CPU threads. The derived
gatefix imatrix is a mechanically augmented copy of that original 256-chunk
file (c8a98474542e2ae7636d7ad0578552398eec3d769b6f9a2981f111a4ba06af12).
Laguna's fused up/gate graph records the shared activation stream under the
up-projection tensor name, so the original imatrix had the up-projection entries
but not all matching gate-projection names. The fix clones:.ffn_up.weight to .ffn_gate.weight.ffn_up_exps.weight to .ffn_gate_exps.weight.ffn_up_shexp.weight to .ffn_gate_shexp.weightoutput.weight and token_embd.weight, which are
explicitly kept at q6_K for these publication quants.laguna-m1-imatrix-calibration-v1.txt:4d478cc0d10da862b0f70040b0c1d948fb88ddac83a5307059e61f02c206622fbuild_laguna_imatrix_corpus.py, with a companion
manifest at laguna-m1-imatrix-calibration-v1.manifest.json. It combines:ubergarm-imatrix-calibration-corpus-v02.txt
(6fd2187f10ab1b742cf85eaf35cf9fa20863a557f68451d6535195774cf1fd0a,
1,699,017 bytes). That baseline is described as combining v5 rc calibration
text with exllamav3 standard calibration data.laguna-m1-support branch, commit
be7d53ceec0c823c07d540e21828d7d187b02d3b, with a clean worktree. The
included source set covers the Laguna converter/runtime path, imatrix code,
quantizer code, CUDA MMQ/MMVQ dispatch, common CLI/chat helpers, and IQK
quantization code.9aa0b10f45e79d7a7bfc2373375eec1890b92c8c1f76a74d9fa4c1b0bc8aeb5c.
The supplement focuses on Laguna M.1/XS.2 metadata, partial RoPE, SWA/global
attention differences, MoE routing and deferred experts, GGUF tensor naming,
CUDA inference, KV-cache quantization, speculative decoding, imatrix
collection, and publication/reproducibility tasks.calibration_datav3.txt from bartowski and calibration_data_v5_rc.txt from
Tristan Dampf/Kalomaze, but did not concatenate them separately because they are
already represented through the ubergarm v02 baseline. The Poolside technical
report text/PDF, private shell history, download logs, local machine notes, and
the earlier ad hoc laguna-m1-code-calib.txt were deliberately excluded.Q2_K, IQ2_M, IQ2_XS, IQ2_S, IQ3_S, Q4_K_M, IQ4_XS, Q5_K_M, Q6_K, and IQ1_S are intended as broad compatibility artifacts once a runner has Laguna GGUF support.IQ1_S and IQ1_S_R4 are experimental very-low-bit artifacts for quality/size probing.IQ2_K_R4, IQ3_K_R4, IQ4_KS_R4, IQ4_K_R4, IQ5_KS_R4, IQ5_K_R4, IQ5_KS, IQ5_K, and IQ6_K are ik_llama.cpp/Spark/Vector-oriented quants selected for direct CUDA MMQ fast paths in ik_llama.cpp.*_R4 quants are not intended as universal llama.cpp artifacts today; use them with ik_llama.cpp or another runner that explicitly supports those tensor types.-ngl values if
the quant is too large for full offload on the validation machine. The validation
logs and results TSVs record the offload setting used where applicable.17 + 25 = 42. All uploaded non-BF16 quant rows above have completed the
same validation harness; BF16 is the reference row used for KL divergence.1llama-quantize \
2 --keep-split \
3 --partial-requant \
4 --imatrix laguna-m1-q8-public-v1-256-gatefix.imatrix \
5 --output-tensor-type q6_K \
6 --token-embedding-type q6_K \
7 Laguna-M.1-BF16-00001-of-00010.gguf \
8 Laguna-M.1-<QUANT>-imatrix-public-v1.gguf \
9 <QUANT> \
10 $(nproc)1llama-perplexity \
2 -m Laguna-M.1-<QUANT>-imatrix-public-v1-00001-of-00010.gguf \
3 -f laguna-m1-imatrix-calibration-v1.txt \
4 -c 4096 \
5 --chunks 64 \
6 -ngl 99 \
7 -b 64 \
8 -ub 16 \
9 --defer-experts \
10 --seed 42 \
11 --kl-divergence \
12 --kl-divergence-base bf16-reference-logprobs.ctx4096.chunks64.bin