Skip to content
important Local AI

Liquid AI releases draft model LFM2.5-VL-DSpark for faster vision-language model inference

clearly official source

Liquid AI has released the draft model LFM2.5-VL-DSpark (280M parameters) for speculative decoding with the LFM2.5-VL-3B model. According to the company, it speeds up decoding by up to 3.13× on-device and up to 2.66× on the H100, with day-one support in llama.cpp, MLX-VLM and SGLang.

Liquid AI has released LFM2.5-VL-DSpark, a small "draft" model designed to speed up inference of the vision-language model LFM2.5-VL-3B via speculative decoding. According to the company, the draft model has approximately 280M parameters and increases the total parameter count of the deployed model by only 8.9%. Architecturally, it shares its principle with the text drafts of the LFM2.5-DSpark series: it captures the hidden states of the target model at several layers and uses them to propose a block of candidate tokens, with both image patches and text tokens projected into a shared representation of the same dimension before these layers.

According to measurements by Liquid AI on six tasks — VQA, text VQA, image captioning, chart VQA, complex reasoning and multi-turn conversation (based on the MMSpec benchmark) — the model achieves a decoding speedup of 2.30× to 3.13× and an overall speedup of 1.56× to 2.62× on-device with MLX on the M5 Max chip. On the M3 Ultra with llama.cpp, the decoding speedup is 1.57× to 2.14× and the overall speedup is 1.30× to 1.77×. On the H100 GPU, the company reports a decoding speedup of up to 2.66× and an overall speedup of 1.64× to 2.27×. The company emphasizes that speculative decoding only speeds up the decoding phase, not image processing by the vision encoder or prefill, so the overall benefit is lower than the decoding speedup alone (Amdahl's law).

According to the company, the draft model has day-one support in llama.cpp, MLX-VLM and SGLang and is available on Hugging Face in Safetensors and GGUF formats. Speculative decoding is described as exact — the target model verifies every proposed token, so greedy output matches running without the draft, with only the speed changing. The model is part of Liquid AI's open LFM2.5 family, which the company says can be downloaded, fine-tuned and deployed without restrictions.

What changed

Why it matters

For developers working with vision-language models, this is a concrete, ready-made tool: simply attach the draft model to an existing LFM2.5-VL-3B in supported frameworks and get a faster response without changing the output, especially on edge devices like Apple silicon, where compute power is more limited than on datacenter GPUs. Companies deploying VLMs in products (e.g. on-device multimodal assistants) can thus reduce inference latency and costs without needing to change their own model or training pipeline.

Release card

LFM2.5-VL-DSpark

Liquid AI

open weights
Specifications
approximately 280M (drafter, increasing the target model LFM2.5-VL-3B by 8.9%)
Inputs
image patches and text tokens projected into a shared representation for speculative decoding
Availability
Available on Hugging Face in Safetensors and GGUF formats, with day-one support in llama.cpp, MLX-VLM and SGLang
Documented measurements
  • MMSpec, MLX na čipu Apple M5 Max zrychlení dekódování 2,30–3,13×, celkové zrychlení 1,56–2,62× Shows how much faster the model generates a response on-device compared to running without the draft model.
  • MMSpec, llama.cpp na čipu Apple M3 Ultra zrychlení dekódování 1,57–2,14×, celkové zrychlení 1,30–1,77× Shows the inference speedup on a different device and with a different tool compared to running without the draft model.
  • MMSpec, GPU H100 zrychlení dekódování až 2,66×, celkové zrychlení 1,64–2,27× Shows the inference speedup on a datacenter GPU compared to running without the draft model.
According to the sources, it is suitable for
  • speeding up the decoding phase of inference for the vision-language model LFM2.5-VL-3B on both edge devices and GPUs
  • deployment without changing the output, since speculative decoding is exact and greedy output matches running without the draft model
  • integration into llama.cpp, MLX-VLM and SGLang with day-one support
Documented limits
  • does not speed up image processing by the vision encoder or the prefill phase, so the overall benefit is lower than the decoding speedup alone
  • on edge devices, prefill makes up a larger share of total latency, which limits the overall effect of the decoding speedup

The card summarizes information from the article and any dated corrections, with a link to the original source. It is not our assessment of the model. It does not yet have a dedicated editorial profile. Model selection and other announcements →

Two audiences, two different impacts

What this means

01

For individuals

Developers deploying LFM2.5-VL-3B can add the DSpark draft model to speed up decoding without changing the output, with support in llama.cpp, MLX-VLM and SGLang.

What to do Try the draft model LFM2.5-VL-DSpark with LFM2.5-VL-3B in llama.cpp, MLX-VLM or SGLang for your own vision-language tasks.
More practical updates →
02

For a business

Companies running vision-language models on edge devices or GPUs can use DSpark to reduce inference latency and compute costs without having to change the target model.

Development
What to decide Consider deploying LFM2.5-VL-3B with the DSpark draft model in products with on-device or GPU vision-language inference to reduce latency.
More business impacts →
Hugging Face inference optimization LFM2.5-VL speculative decoding vision-language modely

Check the original

Event sources

clearly official source · 1 publisher, 0 independent. We count feeds from the same owner only once.

1
Hugging Face Blog primary source · first detected Accelerating vision-language models with LFM2.5-VL-DSpark