Views
No views yet
<|tool_call_start|> / <|tool_call_end|>)ollama pull batiai/lfm2.5-8b:q4| Quant | Size | Recommended For |
|---|---|---|
| Q2_K_S | ~2.8 GB | 8GB Mac, ultra-compact (imatrix) |
| IQ3_XXS | ~3.2 GB | imatrix, smallest K-class footprint |
| Q3_K_M | ~3.9 GB | 8GB+ Mac, balanced |
| IQ4_XS | ~4.3 GB | imatrix, best size/quality |
| Q4_K_M | ~4.9 GB | 16GB Mac (recommended) |
| Q6_K | ~6.5 GB | near-original quality |
Mac note onQ3_K_M: in every model we've benchmarked on Apple Silicon, Q3_K_M generated slower than Q4_K_M despite the smaller file — Granite 4.1 (+27%), Gemma 4 26B (+12%), Qwen3.8‑27B (+18%), Qwen3.6‑27B (+8%), on both M4 Max and M4 mini. Metal's Q3_K path is limited by dequantization compute rather than bandwidth. We have not measured this particular model's Q3/Q4 pair yet, so treat it as a strong prior, not a measurement: if Q4_K_M fits, take it. On CUDA the two are effectively tied, so this applies to Macs only.
| Your Mac RAM | IQ3 | Q2 | Q3 | IQ4 | Q4 | Q6 |
|---|---|---|---|---|---|---|
| 8GB | ✅ | ✅ | ✅ | ✅ | ⚠️ | ❌ |
| 16GB | ✅ | ✅ | ✅ | ✅ | ✅ Recommended | ✅ |
| 24GB+ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |
| BatiAI | Official Liquid GGUF | |
|---|---|---|
| Source | Official Liquid weights | Official |
| Ollama | ✅ batiai/lfm2.5-8b | ❌ HF only |
| Low quants | ✅ IQ3_XXS, Q2_K_S, Q3_K_M | ❌ Q4_0 floor |
| imatrix | ✅ IQ variants calibrated | Standard |
| Tool calling | ✅ Verified | — |
| BatiAI signed | ✅ general.author=BatiAI | — |
lfm2_moe hybrid — 18 LIV conv + 6 GQA layers, 32 experts / 4 active per tokenbench.sh.