GGUF quantized versions of LiquidAI/LFM2.5-2.6B, a high-performance hybrid model designed for on-device deployment, featuring a 128K context window and advanced agentic capabilities.
Model Overview
LFM2.5-2.6B is part of the LFM2.5 family, building on the LFM2 architecture to provide best-in-class performance for its size. It is specifically optimized for agentic workloads, tool use, and long-context workflows, offering competitive performance against models 4x its size.
Key features include:
Agentic Post-Training: Trained using agentic reinforcement learning for improved tool use and instruction following.
Efficient Inference: Designed for high-speed execution on both CPU and GPU.
Reasoning Capabilities: A pure reasoning model that utilizes a <think> tag to reason before answering.
LFM2.5 uses a ChatML-like format. It is a reasoning model that automatically adds a <think> tag when starting an assistant answer to process its logic before providing the final response.
Example format:
<|startoftext|><|im_start|>system
You are a helpful assistant trained by Liquid AI.<|im_end|>
<|im_start|>user
What is C. elegans?<|im_end|>
<|im_start|>assistant
<think>
... reasoning process ...
</think>
C. elegans is a species of small roundworm...<|im_end|>
Tool Calling
LFM2.5 supports Pythonic function calling. It outputs function calls between <|tool_call_start|> and <|tool_call_end|> tokens.
Generation Parameters
Recommended parameters for optimal performance:
Parameter
Value
Temperature
0.1
Top-K
50
Repetition Penalty
1.1
Quantization
These GGUF files were created using llama.cpp tools to enable efficient local deployment on CPUs and GPUs with reduced memory footprints.