Views
No views yet
[!WARNING] This GGUF seems to be broken. This is according to my own testing. Feel free to try it out yourself, and please confirm or deny whether it's broken in the Community tab.
| GGUF Link | Quantization | Description |
|---|---|---|
| Download | Q2_K | Lowest quality |
| Download | Q3_K_S | |
| Download | IQ3_S | Integer quant, preferable over Q3_K_S |
| Download | IQ3_M | Integer quant |
| Download | Q3_K_M | |
| Download | Q3_K_L | |
| Download | IQ4_XS | Integer quant |
| Download | Q4_K_S | Fast with good performance |
| Download | Q4_K_M | Recommended: Perfect mix of speed and performance |
| Download | Q5_K_S | |
| Download | Q5_K_M | |
| Download | Q6_K | Very good quality |
| Download | f16 | Full precision, don't bother; use a quant |
| Property | Value |
|---|---|
| Total Parameters | 854,386,752 (0.85B) |
| Active Parameters | 677,439,552 (0.68B) |
| Architecture | Qwen3.5 Hybrid MoE |
| Experts | 8 routed + 1 shared, top-2 |
| Hidden Size | 1024 |
| Layers | 24 (hybrid: DeltaNet + full attention) |
| Attention | GQA 8Q / 2KV, head_dim=256 |
| Context | 262,144 tokens |
| Vocab | 248,320 |
| Dtype | bfloat16 |
| Component | Source | Strategy |
|---|---|---|
| Embeddings, LM Head | Qwen/Qwen3.5-0.8B | Exact copy |
| Attention (Q/K/V/O, norms) | Qwen/Qwen3.5-0.8B | Exact copy |
| DeltaNet (linear attention) | Qwen/Qwen3.5-0.8B | Exact copy |
| Vision encoder | Qwen/Qwen3.5-0.8B | Exact copy |
| Layer norms | Qwen/Qwen3.5-0.8B | Exact copy |
| Routed experts | Qwen3.5-35B-A3B | Slice 256->8, bilinear resize |
| Shared expert | Qwen3.5-35B-A3B | Bilinear resize |
| Router | Qwen3.5-35B-A3B | Slice + resize |