LFM2 is a family of hybrid models designed for on-device deployment. LFM2-24B-A2B is the largest model in the family, scaling the architecture to 24 billion parameters while keeping inference efficient.
Best-in-class efficiency: A 24B MoE model with only 2B active parameters per token, fitting in 32 GB of RAM for deployment on consumer laptops and desktops.
Fast edge inference: 112 tok/s decode on AMD CPU, 293 tok/s on H100. Fits in 32B GB of RAM with day-one support llama.cpp, vLLM, and SGLang.
Predictable scaling: Quality improves log-linearly from 350M to 24B total parameters, confirming the LFM2 hybrid architecture scales reliably across nearly two orders of magnitude.
image
Find more information about LFM2-24B-A2B in our blog post.
🗒️ Model Details
LFM2-24B-A2B is a general-purpose instruct model (without reasoning traces) with the following features:
Agentic tool use: Native function calling, web search, structured outputs. Ideal as the fast inner-loop model in multi-step agent pipelines.
Offline document summarization and Q&A: Run entirely on consumer hardware for privacy-sensitive workflows (legal, medical, corporate).
Privacy-preserving customer support agent: Deployed on-premise at a company, handles multi-turn support conversations with tool access (database lookups, ticket creation) without data leaving the network.
Local RAG pipelines: Serve as the generation backbone in retrieval-augmented setups on a single machine without GPU servers.
<|startoftext|><|im_start|>system
You are a helpful assistant trained by Liquid AI.<|im_end|>
<|im_start|>user
What is C. elegans?<|im_end|>
<|im_start|>assistant
LFM2-24B-A2B supports function calling as follows:
Function definition: We recommend providing the list of tools as a JSON object in the system prompt. You can also use the tokenizer.apply_chat_template() function with tools.
Function call: By default, LFM2-24B-A2B writes Pythonic function calls (a Python list between <|tool_call_start|> and <|tool_call_end|> special tokens), as the assistant answer. You can override this behavior by asking the model to output JSON function calls in the system prompt.
Function execution: The function call is executed, and the result is returned as a "tool" role.
Final answer: LFM2-24B-A2B interprets the outcome of the function call to address the original user prompt in plain text.
<|startoftext|><|im_start|>system
List of tools: [{"name": "get_candidate_status", "description": "Retrieves the current status of a candidate in the recruitment process", "parameters": {"type": "object", "properties": {"candidate_id": {"type": "string", "description": "Unique identifier for the candidate"}}, "required": ["candidate_id"]}}]<|im_end|>
<|im_start|>user
What is the current status of candidate ID 12345?<|im_end|>
<|im_start|>assistant
<|tool_call_start|>[get_candidate_status(candidate_id="12345")]<|tool_call_end|>Checking the current status of candidate ID 12345.<|im_end|>
<|im_start|>tool
[{"candidate_id": "12345", "status": "Interview Scheduled", "position": "Clinical Research Associate", "date": "2023-11-20"}]<|im_end|>
<|im_start|>assistant
The candidate with ID 12345 is currently in the "Interview Scheduled" stage for the position of Clinical Research Associate, with an interview date set for 2023-11-20.<|im_end|>
🏃 Inference
LFM2-24B-A2B is supported by many inference frameworks. See the Inference documentation for the full list.
If you haven't already, you can install the Transformers.js JavaScript library from NPM using:
npm i @huggingface/transformers
You can then use the model as follows:
js
1import{ pipeline,TextStreamer}from"@huggingface/transformers";23// Create a text generation pipeline4const generator =awaitpipeline(5"text-generation",6"onnx-community/LFM2-24B-A2B-ONNX",7{dtype:"q4f16",device:"webgpu"},8);910// Define the list of messages11const messages =[12{role:"user",content:"What's the capital of France?"},13];1415// Generate a response16const output =awaitgenerator(messages,{17max_new_tokens:512,18do_sample:false,19streamer:newTextStreamer(generator.tokenizer,{20skip_prompt:true,21skip_special_tokens:true,22}),23});24console.log(output[0].generated_text.at(-1).content);
We compared LFM2-24B-A2B against two popular MoE models of similar size: Qwen3-30B-A3B-Instruct-2507 (30.5B total, 3.3B active parameters) and gpt-oss-20b (21B total, 3.6B active parameters). We measured both prefill and decode throughputs with Q4_K_M versions of these models using llama.cpp on AMD Ryzen AI Max+ 395.
image
image
GPU Inference
We also report throughput (total tokens / wall time) achieved with vLLM on a single H100 SXM5 GPU.
image
Contact
For enterprise solutions and edge deployment, contact sales@liquid.ai.
Citation
bibtex
1@article{liquidAI202624B,
2 author = {Liquid AI},
3 title = {LFM2.5-24B-A2B: Scaling Up the LFM2 Architecture},
4 journal = {Liquid AI Blog},
5 year = {2026},
6 note = {www.liquid.ai/blog/},
7}