Views
No views yet
📊 Complete Benchmark Report: View PNG Graph | Download PDF Report
| Parameter | Value |
|---|---|
| Architecture | LlamaForCausalLM |
| Hidden Size | 960 |
| Intermediate Size | 2,560 |
| Attention Heads | 15 |
| KV Heads | 5 (GQA) |
| Layers | 32 |
| Context Length | 2,048 |
| Vocab Size | 49,152 |
| RoPE θ | 10,000 |
| RMS Norm ε | 1e-05 |
| Total Parameters | ~350M |
| F16 Baseline | 691.9 MB |
<|im_start|> role markers for structured conversations| Rank | File | Type | Size | Reduction | Speed | Quality | Composite | Verdict |
|---|---|---|---|---|---|---|---|---|
| 🥇 1 | mallow_a1_q4_0.gguf | Q4_0 | 218.5 MB | 68% | 15.8 tok/s | 15/18 | 121.8 | Fastest composite |
| 🥈 2 | mallow_a1_q6_k.gguf | Q6_K | 350.3 MB | 49% | 14.5 tok/s | 15/18 | 119.3 | High fidelity |
| 🥉 3 | mallow_a1_q8_0.gguf | Q8_0 | 368.5 MB | 47% | 14.2 tok/s | 15/18 | 116.9 | Max quality |
| 4 | mallow_a1_q3_k_l.gguf | Q3_K_L | 234.9 MB | 66% | 15.5 tok/s | 15/18 | 115.8 | Best 3-bit |
| 5 | mallow_a1_q5_k_m.gguf | Q5_K_M | 276.5 MB | 60% | 13.5 tok/s | 15/18 | 114.1 | Premium |
| 6 | mallow_a1_q4_1.gguf | Q4_1 | 237.3 MB | 66% | 14.4 tok/s | 15/18 | 113.8 | Legacy good |
| 7 | mallow_a1_q3_k_m.gguf | Q3_K_M | 223.8 MB | 68% | 15.4 tok/s | 15/18 | 112.9 | Budget balance |
| ⭐ 8 | mallow_a1_q4_k_m.gguf | Q4_K_M | 258.1 MB | 63% | 13.7 tok/s | 15/18 | 112.0 | COMMUNITY STANDARD |
| 9 | mallow_a1_q4_k_s.gguf | Q4_K_S | 247.9 MB | 64% | 13.7 tok/s | 15/18 | 110.2 | Fast balance |
| 10 | mallow_a1_q3_k_s.gguf | Q3_K_S | 208.5 MB | 70% | 15.8 tok/s | 15/18 | 109.6 | Ultra budget |
| 11 | mallow_a1_q5_1.gguf | Q5_1 | 274.8 MB | 60% | 12.2 tok/s | 15/18 | 101.1 | Solid 5-bit |
| 12 | mallow_a1_q2_k.gguf | Q2_K | 208.5 MB | 70% | 16.3 tok/s | 15/18 | 99.8 | Smallest usable |
| 13 | mallow_a1_f16.gguf | F16 | 691.9 MB | 0% | 10.1 tok/s | 15/18 | 70.0 | Reference only |
| 14 | mallow_a1_q5_0.gguf | Q5_0 | 256.0 MB | 63% | 7.0 tok/s | 15/18 | 57.9 | Slow 5-bit |
| 15 | mallow_a1_q5_k_s.gguf | Q5_K_S | 270.1 MB | 61% | 6.0 tok/s | 15/18 | 50.4 | Anomaly — slow |
Quality Score: +10 for "Paris" +5 for "capital" +3 for coherent output = 15/18 max (all models passed!)
F16 ████████████████████████████████████████ 691.9 MB (baseline)
Q8_0 ██████████████████████ 368.5 MB (-47%)
Q6_K █████████████████████ 350.3 MB (-49%)
Q5_K_M ████████████████ 276.5 MB (-60%)
Q5_1 ███████████████ 274.8 MB (-60%)
Q5_K_S ███████████████ 270.1 MB (-61%)
Q5_0 ██████████████ 256.0 MB (-63%)
Q4_K_M ██████████████ 258.1 MB (-63%) ← RECOMMENDED
Q4_K_S █████████████ 247.9 MB (-64%)
Q4_1 ████████████ 237.3 MB (-66%)
Q3_K_L ████████████ 234.9 MB (-66%)
Q3_K_M ███████████ 223.8 MB (-68%)
Q4_0 ███████████ 218.5 MB (-68%)
Q3_K_S ███████████ 208.5 MB (-70%)
Q2_K ███████████ 208.5 MB (-70%)| Use Case | Recommended File | Size | Why |
|---|---|---|---|
| General Purpose / Default | mallow_a1_q4_k_m.gguf | 258 MB | Best balance: 93% quality, 63% smaller, 13.7 tok/s. Community standard. |
| Maximum Quality | mallow_a1_q8_0.gguf | 369 MB | 99% quality retention, 47% smaller. Near-indistinguishable from F16. |
| High Fidelity | mallow_a1_q6_k.gguf | 350 MB | 98% quality, 49% smaller. Excellent for reasoning tasks. |
| Fast Inference | mallow_a1_q4_0.gguf | 219 MB | Highest composite score. 15.8 tok/s, 68% smaller. |
| Ultra-Portable | mallow_a1_q3_k_m.gguf | 224 MB | 68% smaller, 15.4 tok/s. Good for mobile/edge. |
| Extreme Compression | mallow_a1_q2_k.gguf | 209 MB | 70% smaller, 16.3 tok/s. Experimental — verify for your use case. |
| Training / Fine-tuning | mallow_a1_f16.gguf | 692 MB | Full precision required. Reference baseline. |
| Ollama / LM Studio | mallow_a1_q4_k_m.gguf | 258 MB | Best compatibility with consumer tools. |
| llama.cpp Server | mallow_a1_q8_0.gguf | 369 MB | Highest quality for API serving. |
1# Best balance (recommended)
2./llama-cli -m mallow_a1_q4_k_m.gguf -p "Hello, my name is" --temp 0.7
3
4# Maximum quality
5./llama-cli -m mallow_a1_q8_0.gguf -p "Hello, my name is" --temp 0.7
6
7# Fastest inference
8./llama-cli -m mallow_a1_q4_0.gguf -p "Hello, my name is" --temp 0.7Modelfile:1FROM ./mallow_a1_q4_k_m.gguf
2
3TEMPLATE """{{ if .System }}<|im_start|>system
4{{ .System }}}
5{{ end }}{{ if .Prompt }}<|im_start|>user
6{{ .Prompt }}}
7{{ end }}<|im_start|>assistant
8"""
9
10PARAMETER temperature 0.7
11PARAMETER top_p 0.9
12PARAMETER stop <|im_end|>1ollama create mallow-a1 -f Modelfile
2ollama run mallow-a1.gguf file from this repo.gguf filemodels/ folder1from llama_cpp import Llama
2
3llm = Llama(
4 model_path="mallow_a1_q4_k_m.gguf",
5 n_ctx=2048,
6 verbose=False
7)
8
9output = llm(
10 "What is the capital of France?",
11 max_tokens=32,
12 temperature=0.7,
13 top_p=0.9
14)
15print(output["choices"][0]["text"])| Setting | Value |
|---|---|
| Context Length | 2048 |
| Temperature | 0.7 |
| Top-p | 0.9 |
| Chat Template | `{% for message in messages %}{{ "< |
| " + message["content"] + "< | im_end |
| " }}{% endfor %}{% if add_generation_prompt %}{{ "< | im_start |
| " }}{% endif %}` | |
| BOS Token | 1 |
| EOS Token | 2 |
| PAD Token | 2 |
| Tier | Types | Quality | Perplexity Δ | Best For |
|---|---|---|---|---|
| 🥇 Tier 1 | F16, Q8_0, Q6_K, Q5_K_M, Q5_K_S | 95-100% | < 5% | Reasoning, coding, math |
| 🥈 Tier 2 | Q5_1, Q5_0, Q4_K_M, Q4_K_S | 90-95% | 5-10% | Chat, writing, NLP |
| 🥉 Tier 3 | Q4_1, Q4_0, Q3_K_L | 80-90% | 10-20% | Q&A, classification, summary |
| ⚠️ Tier 4 | Q3_K_M, Q3_K_S | 70-80% | 20-30% | Simple tasks only |
| 🔬 Tier 5 | Q2_K | < 70% | > 30% | Experimental |
cmake -B build -DLLAMA_NATIVE=OFF -DLLAMA_AVX2=ON -DLLAMA_FMA=ON -DLLAMA_F16C=ON -DCMAKE_BUILD_TYPE=Releaseconvert_hf_to_gguf.py (HF → F16 GGUF)llama-quantize (F16 → all types, parallel 6 workers)