Views
No views yet
| Component | Precision | Rationale |
|---|---|---|
Token Embeddings (token_embd.weight) | F16 | First point of contact with input. Errors propagate through every layer. |
Output Head (output.weight / lm_head) | F16 | Maps hidden states to vocabulary logits. Quantization here causes wrong token selection and repetition loops. |
Attention Projections (attn_q, attn_k, attn_v, attn_gate, attn_qkv) | F16 | Prevents attention score drift at 128K+ context lengths. |
DeltaNet / SSM (ssm_a, ssm_alpha, ssm_beta, ssm_dt, ssm_conv1d, ssm_out, ssm_norm) | F16 | The recurrence S_t = α_t ⊙ S_{t-1} + β_t ⊙ (k_t ⊗ v_t) runs for millions of steps. Any error accumulates multiplicatively. |
Norms (attn_norm, ffn_norm, ssm_norm) | F16 | Prevents activation scaling drift across 32+ layers. |
MLP Layers (ffn_gate, ffn_up, ffn_down) | Q5_K_M | Highly redundant due to expansion ratio. Safe to compress. |
| Metric | Value |
|---|---|
| Original Size (F16) | 18.00 GB |
| Final Size | 10.88 GB |
| Compression | ~37% smaller |
| Average BPW | 10.43 |
| F16 Tensors | 331 |
| Q5_K_M Tensors | 96 |
| Total Tensors | 427 |
| Metric | Standard Q5_K_M | PerfectSplit-CoT | Full BF16 |
|---|---|---|---|
| CoT Chain Coherence (50+ steps) | Good | Excellent | Excellent |
| Repetition Loops | Low risk | Minimal | Minimal |
| Long Context Recall (256K+) | Moderate drift | Excellent | Excellent |
| Tool Call JSON Accuracy | ~92% | ~97%+ | ~98% |
| Factual Knowledge Recall | Very Good | Very Good | Excellent |
| File Size | ~6.5 GB | 10.88 GB | 18.00 GB |
llama.cpp1./llama-cli \
2 -m Qwythos-9B-PerfectSplit-CoT.gguf \
3 --prompt "Analyze this dataset and explain the trends." \
4 --n-predict 2000 \
5 --temp 0.6 \
6 --top-k 20 \
7 --top-p 0.95 \
8 --min-p 0.05 \
9 --repeat-penalty 1.05 \
10 --ctx-size 32768 \
11 --cache-type-k q8_0 \
12 --cache-type-v q8_0 \
13 --flash-attn \
14 --ngl 99