GGUF quantizations of LiquidAI/LFM2.5-2.6B, a 2.69B-parameter dense hybrid model built for agentic, on-device workloads with a 128K-token context window and native tool calling [web:41][web:39].
Model Details
LFM2.5-2.6B combines 22 double-gated short convolution (LIV) blocks with 8 grouped-query attention (GQA) blocks across 30 total layers, an architecture selected via hardware-in-the-loop search on real edge silicon [web:40][web:44]. It was pre-trained on roughly 34 trillion tokens and post-trained through a four-stage pipeline (SFT, teacher specialization, multi-domain on-policy distillation, and agentic RL) to reliably plan, call tools, and execute multi-step tasks inside agent harnesses [web:38][web:44].
The model uses a ChatML-like chat template with native tool-call tokens (<|tool_call_start|>, <|tool_call_end|>) and a Pythonic tool-call format (function_name(arg="value")) [web:40].
Recommended Quantization
For 8 GB VRAM laptops (e.g. RTX 4060 Laptop), Q4_K_M offers the best balance of quality and footprint (~1.67 GB), leaving headroom for KV cache at long context lengths. For maximum fidelity on the same hardware, Q6_K or Q8_0 still fit comfortably given the model's small base size [web:51].
License
This model inherits the LFM Open License v1.0 from the original LiquidAI/LFM2.5-2.6B release, not a permissive license like Apache 2.0 or MIT — review the terms before commercial deployment [web:50][web:51].
Credits
Original model and architecture by Liquid AI [web:38]. GGUF conversion by NANI-Nithin.