Liquid AI releases draft model LFM2.5-VL-DSpark for faster vision-language model inference
Liquid AI has released the draft model LFM2.5-VL-DSpark (280M parameters) for speculative decoding with the LFM2.5-VL-3B model. According to the company, it speeds up decoding by up to 3.13× on-device and up to 2.66× on the H100, with day-one support in llama.cpp, MLX-VLM and SGLang.
Liquid AI has released LFM2.5-VL-DSpark, a small "draft" model designed to speed up inference of the vision-language model LFM2.5-VL-3B via speculative decoding. According to the company, the draft model has approximately 280M parameters and increases the total parameter count of the deployed model by only 8.9%. Architecturally, it shares its principle with the text drafts of the LFM2.5-DSpark series: it captures the hidden states of the target model at several layers and uses them to propose a block of candidate tokens, with both image patches and text tokens projected into a shared representation of the same dimension before these layers.
According to measurements by Liquid AI on six tasks — VQA, text VQA, image captioning, chart VQA, complex reasoning and multi-turn conversation (based on the MMSpec benchmark) — the model achieves a decoding speedup of 2.30× to 3.13× and an overall speedup of 1.56× to 2.62× on-device with MLX on the M5 Max chip. On the M3 Ultra with llama.cpp, the decoding speedup is 1.57× to 2.14× and the overall speedup is 1.30× to 1.77×. On the H100 GPU, the company reports a decoding speedup of up to 2.66× and an overall speedup of 1.64× to 2.27×. The company emphasizes that speculative decoding only speeds up the decoding phase, not image processing by the vision encoder or prefill, so the overall benefit is lower than the decoding speedup alone (Amdahl's law).
According to the company, the draft model has day-one support in llama.cpp, MLX-VLM and SGLang and is available on Hugging Face in Safetensors and GGUF formats. Speculative decoding is described as exact — the target model verifies every proposed token, so greedy output matches running without the draft, with only the speed changing. The model is part of Liquid AI's open LFM2.5 family, which the company says can be downloaded, fine-tuned and deployed without restrictions.
Why it matters
For developers working with vision-language models, this is a concrete, ready-made tool: simply attach the draft model to an existing LFM2.5-VL-3B in supported frameworks and get a faster response without changing the output, especially on edge devices like Apple silicon, where compute power is more limited than on datacenter GPUs. Companies deploying VLMs in products (e.g. on-device multimodal assistants) can thus reduce inference latency and costs without needing to change their own model or training pipeline.
Release card
LFM2.5-VL-DSpark
Liquid AI
- Specifications
- approximately 280M (drafter, increasing the target model LFM2.5-VL-3B by 8.9%)
- Inputs
- image patches and text tokens projected into a shared representation for speculative decoding
- Availability
- Available on Hugging Face in Safetensors and GGUF formats, with day-one support in llama.cpp, MLX-VLM and SGLang
- MMSpec, MLX na čipu Apple M5 Max zrychlení dekódování 2,30–3,13×, celkové zrychlení 1,56–2,62× Shows how much faster the model generates a response on-device compared to running without the draft model.
- MMSpec, llama.cpp na čipu Apple M3 Ultra zrychlení dekódování 1,57–2,14×, celkové zrychlení 1,30–1,77× Shows the inference speedup on a different device and with a different tool compared to running without the draft model.
- MMSpec, GPU H100 zrychlení dekódování až 2,66×, celkové zrychlení 1,64–2,27× Shows the inference speedup on a datacenter GPU compared to running without the draft model.
- speeding up the decoding phase of inference for the vision-language model LFM2.5-VL-3B on both edge devices and GPUs
- deployment without changing the output, since speculative decoding is exact and greedy output matches running without the draft model
- integration into llama.cpp, MLX-VLM and SGLang with day-one support
- does not speed up image processing by the vision encoder or the prefill phase, so the overall benefit is lower than the decoding speedup alone
- on edge devices, prefill makes up a larger share of total latency, which limits the overall effect of the decoding speedup
The card summarizes information from the article and any dated corrections, with a link to the original source. It is not our assessment of the model. It does not yet have a dedicated editorial profile. Model selection and other announcements →
Two audiences, two different impacts
What this means
For individuals
Developers deploying LFM2.5-VL-3B can add the DSpark draft model to speed up decoding without changing the output, with support in llama.cpp, MLX-VLM and SGLang.
For a business
Companies running vision-language models on edge devices or GPUs can use DSpark to reduce inference latency and compute costs without having to change the target model.
DevelopmentCheck the original
Event sources
clearly official source · 1 publisher, 0 independent. We count feeds from the same owner only once.