LFM2.5-VL-3B model released with improved screen reading and function calling for edge devices
The vision-language model LFM2.5-VL-3B (3.1 billion parameters) has been released for edge devices. According to the developers, it improves screen reading, working with multiple images and function calling, has open weights and runs directly on the device.
The vision-language model LFM2.5-VL-3B has been released with 3.1 billion parameters, designed for deployment on edge devices. According to the developers, it brings four main improvements over the previous version: understanding screens and user interfaces on different types of devices, more accurate grounding and object detection based on natural language, better reasoning over multiple images simultaneously and significantly stronger function calling in both text and visual-text tasks.
The model combines the SigLIP2 visual encoder with 400 million parameters (the NaFlex variant) with the LFM2.5-2.6B language backbone. Pretraining used approximately 34 trillion tokens, with four times the volume of visual data compared with the previous version. The tokenizer vocabulary was expanded to 128K entries to support languages that use non-Latin scripts. Post-training took place in two phases: supervised fine-tuning with knowledge distillation from a larger teacher model and so-called Antidoom training, followed by reinforcement learning with multiple reward signals.
According to the published results, the model achieves an average score of 69.4 across a set of visual benchmarks (including MMStar, ChartQA, DocVQA, RefCOCO, BLINK, ScreenSpot), matching the result of the larger model InternVL 3.5 4B, while Qwen3.5-4B (also a larger model) achieved a higher average of 70.1. In text tasks focused on tool use and function calling, LFM2.5-VL-3B is comparable to Gemma-4-E2B and Qwen3.5-2B. The model has had support in llama.cpp, MLX, vLLM, SGLang and ONNX since release. According to the published measurements, it handles 228 tokens per second on the M5 Max chip, 116 tokens per second on Ryzen AI Max+ 395, fits into roughly 3 GB of memory and runs at 20 tokens per second on the Galaxy S26 Ultra phone.
You can find details in the source article.
Why it matters
A small open vision-language model with strong function calling and screen reading capabilities makes it possible to build agents and assistants that run directly on a laptop or mobile phone without calling a cloud API — this reduces latency, inference costs and reliance on external infrastructure for tasks such as UI automation, working with documents or mobile assistants.
Release card
LFM2.5-VL-3B
Liquid AI
- Specifications
- 3.1 billion parameters (3.1B)
- Inputs
- input: text and images (support for multiple images at once), output: text
- Availability
- The model is available for deployment on edge devices, with support in llama.cpp, MLX, vLLM, SGLang and ONNX from the day of release.
- MMStar 63,3 Score on the MMStar benchmark for general visual understanding.
- Průměr napříč vizuálními benchmarky 69,4 The overall average score of the model across a set of visual tasks (general understanding, documents, grounding, GUI and others).
- ChartQA (test) 81,3 Score on the task of reading and interpreting charts.
- ScreenSpot-v2 Mobile 81,2 Score on the task of locating elements on a mobile screen based on a text description.
- BFCL V4 32,5 Score on the benchmark for function calling.
- Reading and understanding the content of screens and user interfaces on devices
- Function calling and tool calling in both text and visual-text tasks
- Understanding documents, charts and multiple images simultaneously, running directly on the device (edge)
The card summarizes information from the article and any dated corrections, with a link to the original source. It is not our assessment of the model. It does not yet have a dedicated editorial profile. Model selection and other announcements →
Two audiences, two different impacts
What this means
For individuals
Developers can deploy an open model with 3.1 billion parameters directly on a laptop or mobile phone for tasks such as screen reading, working with multiple images or function calling without relying on a cloud API.
For a business
Companies developing edge or mobile applications with AI gain a smaller open model with strong function calling and screen reading capabilities, enabling it to run directly on the device instead of using paid cloud inference.
DevelopmentCheck the original
Event sources
clearly official source · 1 publisher, 0 independent. We count feeds from the same owner only once.