Views
No views yet
| Mean KLD (abliterated BF16 vs original BF16) | 0.0151 |
| same top-1 token | 95.04% |
| tier | size | mean KLD | 99% KLD | 99.9% KLD | PPL ratio | same top-1 | active bpw | file bpw | overall |
|---|---|---|---|---|---|---|---|---|---|
MTP-UD-Q6_K | 30.21 GB | 0.0222 | 0.226 | 0.695 | 0.9960 | 94.05% | 8.063 | 6.804 | 96.7 |
MTP-UD-Q5_K_S | 25.84 GB | 0.0261 | 0.269 | 0.899 | 0.9871 | 93.51% | 7.693 | 5.820 | 96.3 |
TIEL_Calibrated-25G-ICE | 24.85 GB | 0.0276 | 0.269 | 1.061 | 0.9847 | 93.33% | 7.686 | 5.597 | 96.1 |
MTP-25G-ICE | 24.85 GB | 7.686 | 5.597 | ||||||
TIEL_Calibrated-23G-ICE | 22.84 GB | 0.0321 | 0.326 | 1.173 | 0.9907 | 92.89% | 7.523 | 5.143 | 95.7 |
MTP-23G-ICE | 22.84 GB | 7.523 | 5.143 | ||||||
TIEL_Calibrated-21G-ICE | 20.85 GB | 0.0392 | 0.421 | 1.606 | 0.9997 | 92.27% | 7.357 | 4.695 | 95.0 |
MTP-21G-ICE | 20.85 GB | 7.357 | 4.695 | ||||||
TIEL_Calibrated-19G-ICE | 18.82 GB | 0.0576 | 0.602 | 2.123 | 1.0015 | 90.42% | 7.192 | 4.240 | 93.3 |
MTP-19G-ICE | 18.82 GB | 7.192 | 4.240 | ||||||
MTP-UD-IQ4_XS | 18.68 GB | 0.0706 | 0.695 | 2.645 | 1.0502 | 89.40% | 6.762 | 4.209 | 92.2 |
MTP-APEX-I-Compact-v2D-lite | 17.57 GB | 0.0925 | 0.893 | 3.141 | 1.0169 | 87.97% | 5.228 | 3.956 | 90.5 |
MTP-APEX-I-Mini-v2D-lite | 14.37 GB | 0.2546 | 2.410 | 5.578 | 1.2129 | 80.70% | 4.180 | 3.208 | 80.0 |
The fourICErows were rebuilt on 2026-08-24 and the struck-through values are what they replaced. The new figures are measured at 64 chunks, the rest of the table at 16; a higher chunk count tightens the error bar without moving the mean. What changed, why, and the full measurement set:MEASUREMENTS-ICE-rebuild.md.
The fourTIEL_Calibratedrows were added on 2026-08-30. Identical builds to theirMTP-*-ICEcounterparts — same abliterated trunk, same norm-fixed MTPv2 head, same ICE-base recipes, same byte budgets — but built with the importance matrix and embedded chat template from peculiar-ragdoll/Tiel-Coder-35B-A3B-GGUF, with credit to that repo for both. Its imatrix is 3000 x 512 = 1,536,000 tokens against the 573 x 512 = 293,376 behind the others. Also measured at 64 chunks.
0.70/(1+meanKLD) + 0.30*sameTop1. All KLD columns are measured
against the abliterated BF16, i.e. they isolate what the quantization costs.Read the tail columns with care.99.9% KLDis the ~33rd-worst token out of 32,768 — an extreme order statistic with large sampling variance, so it inverts between adjacent tiers without that meaning anything.99% KLDrests on ~328 tokens and orders all nine tiers monotonically;mean KLDuses all 32,768 and separates the closest pair by 4.3 sigma. Rank on mean KLD; treat the tail columns as shape, not order.
ICE tiers have since been rebuilt; their current numbers
are in the tier table above and in
MEASUREMENTS-ICE-rebuild.md.| tier | KLD abl | KLD clean | Δ | top-1 abl | top-1 clean | Δ | PPL ratio abl | PPL ratio clean | Δ |
|---|---|---|---|---|---|---|---|---|---|
MTP-UD-Q6_K | 0.0222 | 0.0221 | +0.6% | 94.05% | 93.85% | +0.20 pp | 0.9960 | 0.9957 | +0.0003 |
MTP-UD-Q5_K_S | 0.0261 | 0.0272 | -4.2% | 93.51% | 93.51% | +0.01 pp | 0.9871 | 0.9862 | +0.0009 |
MTP-25G-ICE | 0.0293 | 0.0303 | -3.4% | 93.44% | 93.16% | +0.28 pp | 0.9856 | 0.9814 | +0.0042 |
MTP-23G-ICE | 0.0345 | 0.0361 | -4.4% | 92.39% | 92.65% | -0.26 pp | 0.9902 | 0.9885 | +0.0018 |
MTP-21G-ICE | 0.0398 | 0.0412 | -3.4% | 92.06% | 92.03% | +0.03 pp | 0.9943 | 0.9924 | +0.0019 |
MTP-19G-ICE | 0.0612 | 0.0608 | +0.5% | 90.02% | 90.32% | -0.30 pp | 1.0025 | 1.0030 | -0.0005 |
MTP-UD-IQ4_XS | 0.0706 | 0.0723 | -2.4% | 89.40% | 89.46% | -0.06 pp | 1.0502 | 1.0526 | -0.0024 |
MTP-APEX-I-Compact-v2D-lite | 0.0925 | 0.0954 | -3.0% | 87.97% | 87.83% | +0.14 pp | 1.0169 | 1.0101 | +0.0068 |
MTP-APEX-I-Mini-v2D-lite | 0.2546 | 0.2608 | -2.4% | 80.70% | 80.49% | +0.21 pp | 1.2129 | 1.2281 | -0.0151 |
ffn_down narrows its dynamic range.KLD_RESULTS.txt):| family | what it is |
|---|---|
| UD-* | Unsloth Dynamic 2.0 maps, replayed 1:1. Pins attention, the shared expert and token_embd at Q8_0 at every size and moves only the routed experts. |
| ICE-* | Bits allocated by how far a quantization error travels, not by activation magnitude. Named by target size. |
| APEX-I-*-v2D-lite | mudler's APEX maps plus one extra step on attn_k/attn_v in the ten full-attention blocks and on the output head. |
| comparison | result |
|---|---|
| 23G-ICE vs UD-Q4_K_XL (same size) | -5.0% KLD |
| 23G-ICE vs APEX-I-Quality | -13.0% KLD and 0.61 GB smaller |
| 25G-ICE vs APEX-I-Balanced | -12.2% KLD and 1.15 GB smaller |
UD-Q4_K_XL,APEX-I-Quality-v2D-lite and APEX-I-Balanced-v2D-lite are each already covered by
an ICE tier that is both smaller and closer to BF16, so rebuilding them would add size
without adding a quality point. Full derivation, the refuted ffn_down rule
and the measured convexity bound are in the
original Ornith-1.5 card.The 9-Tier Standard
From this release onward these nine recipes are the standard ladder:UD-Q6_K·UD-Q5_K_S·25G-ICE·23G-ICE·21G-ICE·19G-ICE·UD-IQ4_XS·APEX-I-Compact-v2D-lite·APEX-I-Mini-v2D-liteThey are the measured Pareto frontier of a 12-tier sweep on this architecture: every dropped tier is beaten on both size and KLD by one that ships. Reference measurements and methodology: Ornith-1.5-35B-MTP-UD-APEX-GGUF.
PocketAiHub/Ornith-1.5-35B-A3B-Abliterated-GGUF is faithful at the tensor level —
every tensor except ffn_down is byte-identical to ornith-ai's BF16. Abliteration is
confined to ffn_down (both routed and shared experts), layers 15-39, at
~1.6-1.9e-02 L1-relative. Routers, attention, ffn_gate, ffn_up and layers 0-14
are untouched.blk.40.* tensors from ornith-ai's BF16;
block_count 40 -> 41, nextn_predict_layers added.tokenizer.ggml.add_bos_token restored to False — the source omits the key
entirely, so llama.cpp falls back to its own default and tokenises differently
from the original. Left unfixed this also invalidates any KLD against the original.tokenizer.chat_template restored — the source ships a 7536-byte copy with
the multi-system-message merge block removed; the original is 7828 bytes.Ornith-1.5-35B-A3B-imatrix.gguf, reused unmodified. Justified by
measurement, not assumption: the routers are byte-identical between the original
and the abliterated model, so the same experts fire and the per-channel statistics
still apply.unsloth/Ornith-1.0-35B-GGUF (Unsloth Dynamic 2.0).attn_k/attn_v on the ten
full-attention blocks and the output head each raised one step.attn_k/attn_v on the ICE tiers).
It is grafted from the original model, so the draft head is not abliterated while
the trunk is — worth knowing if you rely on the refusal behaviour under drafting:
the head only proposes, the abliterated trunk verifies, so accepted tokens are always
the trunk's.23G-ICE, --spec-type draft-mtp, text-only.| prompt set | acceptance |
|---|---|
code-novel | 98.45% |
structured | 98.29% |
copy-edit | 97.30% |
prose-novel | 91.57% |
gate_spec_bench.json. Acceptance depends on the prompt mix — compare only against numbers taken on the same harness.1llama-server -m Ornith-1.5-35B-A3B-Abliterated-MTPv2-21G-ICE.gguf \
2 --mmproj mmproj-Ornith-1.5-35B-A3B-Abliterated-F16.gguf \
3 -c 8192 -fa on --jinja \
4 --spec-type draft-mtp --spec-draft-n-max 1 --spec-draft-n-min 0 --spec-draft-p-min 0.75ornith-ai uploaded a newly trained native MTP head to the official
Ornith-1.5-35B-A3B repo, so this
repo is updated with it.MTPv2-* files are the same nine recipes as the previous ladder — same
abliterated trunk, same imatrix, same per-tensor rule files, same pinned llama.cpp build.
Only blk.40 is re-grafted. Every tensor outside blk.40, and the whole metadata block,
is byte-identical to its MTPv1/ counterpart — checked tensor by tensor before anything
was uploaded.ICE tiers were later rebuilt again, on 2026-08-24, with a revised expert
placement and a corrected blk.40 norm conversion. Their rows in the tier table show
both the old and the new numbers.blk.40 tensor | MTPv1 | MTPv2 |
|---|---|---|
nextn.shared_head_norm.weight | 1.02281 ± 0.00017 | 2.92531 ± 0.29643 |
attn_norm.weight | 1.00006 ± 0.00302 | 0.90492 ± 0.15616 |
nextn.hnorm.weight | 1.01571 ± 0.00558 | 0.49370 ± 0.09479 |
attn_q_norm.weight | 1.00390 ± 0.00680 | 1.76704 ± 0.33864 |
MTPv1/ — same files, same
digests, nothing deleted.--model-draft sidecar instead
of an embedded head:| file | size |
|---|---|
mtp-Ornith-1.5-35B-A3B-Abliterated-MTPv2-BF16.gguf | 3.74 GB |
mtp-Ornith-1.5-35B-A3B-Abliterated-MTPv2-Q8_0.gguf | 1.99 GB |
mtp-Ornith-1.5-35B-A3B-Abliterated-MTPv2-Q4_K_M.gguf | 1.26 GB |
MTPv2-* files.mmproj-Ornith-1.5-35B-A3B-Abliterated-F16.gguf
(0.90 GB). Without it the model is blind. Note --mmproj force-disables ctx_shift
and cache_reuse.sha256sums-MTPv2.txt and MANIFEST-MTPv2.txt for the current ladder,
MEASUREMENTS-ICE-rebuild.md for the 2026-08-24 ICE rebuild, and the MTPv1
originals under MTPv1/ (KLD_RESULTS.txt, gate_spec_bench.json,
sha256sums.txt, MANIFEST.txt).