Views
No views yet
qwen3_5_moe, ~35B total / A3B active, vision-capable), produced by
MagicQuant's measured evolutionary search: tensors are grouped by role,
candidate per-group scheme assignments are rendered and perplexity-measured
rather than predicted, and a winner is selected per size band. imatrix-
calibrated (510 tensors), KL-blended (weight 0.1), 15 measured candidates over
3 rounds.ed17991, Foundry 2f99202.wiki.test.raw. BF16 baseline: 66.19 GiB, PPL 7.9691.| file | size | ratio vs BF16 | PPL | vs baseline |
|---|---|---|---|---|
Ornith-1.5-35B-A3B-Q4_K_M.gguf | 19.58 GiB | 0.30× | 7.9530 | −0.20% |
Ornith-1.5-35B-A3B-Q5_K_M.gguf | 22.88 GiB | 0.35× | 7.9951 | +0.33% |
Ornith-1.5-35B-A3B-Q6_K.gguf | 28.84 GiB | 0.44× | 7.9792 | +0.13% |
mmproj-Ornith-1.5-35B-A3B-f16.gguf | 0.86 GiB | — | — | vision projector |
Q4, Q5 and Q6 are indistinguishable from BF16 and from each other at this resolution. Separating them would need a sharper instrument (paired KL divergence, roughly 100× the resolution), which was not run.
Q4_K_M. It is 0.30× the BF16 size at a cost this instrument
cannot separate from zero. The larger tiers are published for anyone who wants
headroom, not because they are measurably better here.mmproj file alongside the model.tokenizer.ggml.token_type) is INT32 per spec — these files
load on current mainline llama.cpp builds.qwen3_5_moe. A build from mid-2026 or newer is recommended.