Views
No views yet
LFM2.5-2.6B-<quant>.gguf scheme and is optimized for different trade‑offs between size and inference speed.token_embd.weight layer was not quantized in any of these files.| Quant | File Name | Size (GB) |
|---|---|---|
| F16 | LFM2.5-2.6B-F16.gguf | 5.27 |
| Q8_0 | LFM2.5-2.6B-Q8_0.gguf | 2.80 |
| Q6_K | LFM2.5-2.6B-Q6_K.gguf | 2.16 |
| Q5_K_M | LFM2.5-2.6B-Q5_K_M.gguf | 1.89 |
| Q4_K_M | LFM2.5-2.6B-Q4_K_M.gguf | 1.63 |
| Q4_K_S | LFM2.5-2.6B-Q4_K_S.gguf | 1.56 |
| IQ4_NL | LFM2.5-2.6B-IQ4_NL.gguf | 1.55 |
| IQ4_XS | LFM2.5-2.6B-IQ4_XS.gguf | 1.48 |
| Q3_K_S | LFM2.5-2.6B-Q3_K_S.gguf | 1.24 |
| IQ3_XS | LFM2.5-2.6B-IQ3_XS.gguf | 1.19 |
| Q2_K | LFM2.5-2.6B-Q2_K.gguf | 1.06 |
imatrix quantization tool. This method provides high‑quality compression while preserving model accuracy.token_embd.weight layer remains in full FP16 (or the original precision) and was not quantized. This ensures that embedding look‑ups remain accurate during inference.llama.cpp or auto-gptq):1# Example with llama.cpp
2./llama-server -m "./LFM2.5-2.6B-Q4_K_M.gguf" -c 128000