APEX (Adaptive Precision for EXpert Models) quantizations of Ornith-1.0-35B, an open-source coding MoE model by DeepReinforce (MIT license, based on Qwen 3.5 architecture).
These quants were produced using the apex-quant toolchain. APEX is a MoE-aware mixed-precision quantization strategy that classifies tensors by role (routed expert, shared expert, attention) and applies a layer-wise precision gradient — edge layers get higher precision, middle layers more aggressive compression.
Files
Each profile comes in two variants:
Base — the quantized model standalone
-MTP — includes the bundled MTP (multi-token prediction) head, quantized to Q8_0 (near-lossless), for self-speculative decoding via --spec-type draft-mtp. Requires a recent llama.cpp build with MTP support.
I-variants were calibrated with a diverse importance matrix (chat, code, reasoning, tool-calling, multilingual) for improved downstream accuracy.
File
Profile
Size
Best For
ornith-1.0-35b-APEX-I-Mini.gguf
I-Mini
14 GB
Smallest viable, fastest inference
ornith-1.0-35b-APEX-I-Mini-MTP.gguf
I-Mini + MTP
14 GB
Smallest viable + self-spec
ornith-1.0-35b-APEX-Compact.gguf
Compact
17 GB
Consumer GPUs, general purpose
ornith-1.0-35b-APEX-Compact-MTP.gguf
Compact + MTP
17 GB
Consumer GPUs + self-spec
ornith-1.0-35b-APEX-I-Compact.gguf 🏆
I-Compact
17 GB
Consumer GPUs, best quality at this size
ornith-1.0-35b-APEX-I-Compact-MTP.gguf 🏆
I-Compact + MTP
17 GB
Consumer GPUs, best quality + self-spec
ornith-1.0-35b-APEX-Quality.gguf
Quality
22 GB
Highest quality standard
ornith-1.0-35b-APEX-Quality-MTP.gguf
Quality + MTP
23 GB
Highest quality + self-spec
ornith-1.0-35b-APEX-I-Quality.gguf
I-Quality
22 GB
Highest quality with imatrix
ornith-1.0-35b-APEX-I-Quality-MTP.gguf
I-Quality + MTP
23 GB
Highest quality + imatrix + self-spec
ornith-1.0-35b-APEX-Balanced.gguf
Balanced
24 GB
General purpose, best trade-off
ornith-1.0-35b-APEX-Balanced-MTP.gguf
Balanced + MTP
25 GB
General purpose + self-spec
ornith-1.0-35b-APEX-I-Balanced.gguf 🏆
I-Balanced
24 GB
Best overall — lowest KL divergence
ornith-1.0-35b-APEX-I-Balanced-MTP.gguf 🏆
I-Balanced + MTP
25 GB
Best overall + self-spec
Profile Precision Breakdown
APEX applies a layer-wise precision gradient to MoE expert weights. I-variants additionally use a diverse imatrix (chat, code, reasoning, tool-calling) that improves downstream accuracy and lowers KL divergence.
Profile
Edge (blk 0-4, 35-39)
Near-Edge (blk 5-9, 30-34)
Middle (blk 10-29)
Shared Expert
Attention
Quality
Q6_K
Q5_K
IQ4_XS
Q8_0
Q6_K
Balanced
Q6_K
Q5_K
Q5_K
Q8_0
Q6_K
Compact
Q4_K
Q3_K
Q3_K
Q6_K
Q4_K
Mini
Q3_K_M
Q3_K_M
IQ2_S
Q4_K
Q3_K_M
Quality and Mini use a 3-tier gradient. Balanced and Compact use a simpler 2-tier gradient (edge vs. middle) — their "Near-Edge" and "Middle" columns are the same precision.
MTP Head
The bundled MTP head (blk.40.* including the nextn.* projection + norms) is quantized to Q8_0 (near-lossless) for high draft accuracy. Enable with: