LFM2.5-VL-3B is a multimodal variant of LFM2.5, a family of hybrid models designed for on-device deployment. It builds on LFM2-VL-3B with further mid- and post-training. LFM2.5-VL-3B can process both text and images, and uses the LFM2.5-2.6B language model as its backbone, combined with a SigLIP2 NaFlex vision encoder.
Better grounding: Improved grounding and object detection with natural language queries.
Better OCR: Full page OCR with layout annotation. See layout annotation format for more information.
Efficient inference: 228 tok/s on an Apple M5 Max and 116 tok/s on an AMD Ryzen AI Max+ 395, in under 3.3 GB of memory.
Find more information about LFM2.5-VL-3B in our release post.
lfm2_5_vl_3b_task_group_averages
[!NOTE]
💻 Demos: Try LFM2.5-VL-3B's vision understanding capabilities in a Hugging Face space without any setup:
Vision-capable chat in your browser: allows you to upload images or use the webcam to capture stills and let the model interact with them.
Quantized ONNX exports for cross-platform deployment. Enables hardware-accelerated inference across diverse environments (cloud, edge, mobile). See the demo.
We recommend using it for single-turn, high-throughput, low-latency tasks; for example, for near-realtime object detection in automotive applications, batch processing scanned documents with OCR with layout information for turning PDFs into searchable text, or for on-device translation of menus and road signs into your native language.
It is not recommended for long-context, reasoning-intensive tasks, such as visual web design, or answering highly technical questions about blueprints.
<|startoftext|><|im_start|>system
You are a helpful assistant trained by Liquid AI.<|im_end|>
<|im_start|>user
What species is in this picture?<image><|im_end|>
<|im_start|>assistant
[!TIP]
Note: The apply_chat_template() method automatically inserts the <image> tag for each image in your message. Do not include <image> in your message content.
Inference
LFM2.5-VL is supported by many inference frameworks. See the Inference documentation for the full list.
LFM2.5-VL-3B supports function calling in four steps:
Function definition: Provide the list of tools as a JSON object in the system prompt, or use tokenizer.apply_chat_template() with tools=....
Function call: By default, LFM2.5 writes Pythonic function calls (a Python list between <|tool_call_start|> and <|tool_call_end|> special tokens), as the assistant answer.
Function execution: Execute the call and return the result with the tool role.
Final answer: LFM2.5 interprets the tool output and returns a plain-text answer addressing the original prompt.
<|startoftext|><|im_start|>system
List of tools: [{"name": "get_candidate_status", "description": "Retrieves the current status of a candidate in the recruitment process", "parameters": {"type": "object", "properties": {"candidate_id": {"type": "string", "description": "Unique identifier for the candidate"}}, "required": ["candidate_id"]}}]<|im_end|>
<|im_start|>user
What is the current status of candidate ID 12345?<|im_end|>
<|im_start|>assistant
<|tool_call_start|>[get_candidate_status(candidate_id="12345")]<|tool_call_end|>Checking the current status of candidate ID 12345.<|im_end|>
<|im_start|>tool
[{"candidate_id": "12345", "status": "Interview Scheduled", "position": "Clinical Research Associate", "date": "2023-11-20"}]<|im_end|>
<|im_start|>assistant
The candidate with ID 12345 is currently in the "Interview Scheduled" stage for the position of Clinical Research Associate, with an interview date set for 2023-11-20.<|im_end|>
Layout Annotation Format
LFM2.5-VL-3B can do OCR with layout annotation. The layout annotation is a list of regions, each with a label, bounding box, and content. The format is:
[xmin, ymin, xmax, ymax] are normalized integer coordinates in [0, 1000], same as our grounding format.
<content> is the region's content:
plain text for text regions
LaTeX for equations
OTSL (Optimized Table Structure Language, introduced here by IBM) for tables
a short description for images and charts
There will be a blank line between each region.
To prompt the model to generate this structured output, use a system or user prompt that includes this:
text
1Parse this document into its layout regions. The pages are provided as images in reading order. For every region, in reading order across all pages, output a header line immediately followed by the region's content:
23image_index=<n> <label> [xmin, ymin, xmax, ymax]
4<content>
56where:
7- image_index is the zero-based index of the page image the region appears on (0 for the first image, 1 for the second, and so on)
8- <label> is one of these layout labels: text, title, list, table, table_caption, table_footnote, image, image_block, image_caption, image_footnote, chart, equation, formula_number, code, code_caption, algorithm, aside_text, ref_text, phonetic, page_header, page_footer, page_number, page_footnote
9- [xmin, ymin, xmax, ymax] are normalized integer coordinates in [0, 1000]
10- <content> is the region's content: plain text for text regions, LaTeX for equations, OTSL for tables, and a short description for images and charts
1112Separate each region block with one blank line. Return only the parsed regions.
Note that the layout annotation format is still experimental: it may change, may be unreliable, and may not be trivial to parse. We encourage users to try it out and provide feedback!
Fine-Tuning
We recommend fine-tuning LFM2.5-VL models for your specific use case to achieve the best results.
LFM2.5-VL-3B decodes 228 tokens/s on an Apple M5 Max and 116 tokens/s on an AMD Ryzen AI Max+ 395, and fits in about 3 GB of memory. It even reaches 20 tokens/s on a Galaxy S26 Ultra, so you can run it fully on-device.
lfm2_5_vl_3b_on-device_inference_TTFT
GPU Inference
On a single NVIDIA H100 with vLLM, LFM2.5-VL-3B reaches the highest output throughput of any model we tested, about 11K tokens per second at high concurrency, or nearly 1B tokens per day.
lfm2_5_vl_3b_throughput
Because it answers directly instead of reasoning, LFM2.5-VL-3B is quick to first token on a single H100, reaching about 34 ms on a 5-frame clip.
If you are interested in custom solutions with edge deployment, please contact our sales team.
Citation
@article{liquidAI2026VL3B,
author = {Liquid AI},
title = {LFM2.5-VL-3B: A Better and Faster Vision-Language Model for the Edge},
journal = {Liquid AI Blog},
year = {2026},
note = {www.liquid.ai/blog/lfm2-5-vl-3b},
}