Views
No views yet
llama.cpp build 8044 (91ea5d67f).IQ4_XS was quantized with an imatrix generated on 19th-century public-domain English text.5504) not divisible by 256, so llama.cpp applied fallback quantization to 22 tensors for K/IQ quant types.| Detail | Value |
|---|---|
| Model Architecture | LlamaForCausalLM (decoder-only transformer) |
| Parameter Count | ~1.22B |
| Training Type | Trained from scratch (random initialization) |
| Tokenizer | Custom BPE, vocab size 32,000 |
| Sequence Length | 2048 |
| Attention Type | Grouped Query Attention (16 Q heads / 8 KV heads) |
| Hidden Size | 2048 |
| Intermediate Size | 5504 |
| Layers | 22 |
| Link | Type | Size/GB | Notes |
|---|---|---|---|
| GGUF | Q2_K | 0.5 | smallest |
| GGUF | Q3_K_S | 0.6 | low VRAM |
| GGUF | Q3_K_M | 0.6 | balanced low size |
| GGUF | Q3_K_L | 0.6 | better than Q3_K_M |
| GGUF | IQ4_XS | 0.6 | imatrix, recommended at this size |
| GGUF | Q4_K_S | 0.7 | fast, recommended |
| GGUF | Q4_K_M | 0.7 | fast, recommended |
| GGUF | Q5_K_S | 0.8 | higher quality |
| GGUF | Q5_K_M | 0.9 | higher quality |
| GGUF | Q6_K | 1.0 | very good quality |
| GGUF | Q8_0 | 1.2 | fast, best quality |
| GGUF | f16 | 2.3 | 16 bpw, overkill |