Skip to content
important New models

Liquid AI releases weights for d1-3B and d1-omni-600M for local decision-making

clearly official source

Liquid AI has released the open weights of d1-3B and d1-omni-600M. The models return decisions in a single pass without generating tokens. They support text with images, and the experimental d1-omni-600M also supports text with audio.

Liquid AI has made the open weights of the decision models d1-3B and d1-omni-600M available on Hugging Face for local use. d1-3B accepts text and images, while d1-omni-600M accepts a combination of text and an image or text and audio. According to the company, the models return an answer in a single pass through the network without generating tokens. d1-omni-600M is an early research version that is still under development.

For d1-3B, the company reports a response time for a single question of 8 ms on NVIDIA GeForce RTX 4090, 16 ms on NVIDIA Jetson AGX Thor, 26 ms on NVIDIA Jetson AGX Orin 64 GB and 50 ms on NVIDIA Jetson Orin Nano. Each measurement used a single request. Longer inputs have significantly higher latency: on NVIDIA Jetson Orin Nano, an input of 3 400 tokens took 1 640 ms, and processing a 384 px image took 202 ms. The company has not published speed figures for d1-omni-600M.

According to the company, d1-3B achieved a score of 48.57 on the public portion of Decision Index v0.2.1, and d1-omni-600M scored 15.95. Both articles report averages across seven text benchmarks of 82.9 and 78.4, respectively, but some individual results differ. For example, for d1-3B on SQuAD 2.0, the Liquid AI blog reports 85.3, while the article on Hugging Face reports 83.3. The published results do not include a separate evaluation of decision-making based on images and audio.

The models have had llama.cpp support since release. The article on Hugging Face also includes an example using the transformers library, version 5.14 or later, which simultaneously identifies a refund request, the appropriate team and the urgency within a single customer request. Demos are available to try in the System One Arcade Space on Hugging Face.

What changed

Why it matters

Developers of applications that respond to camera images can evaluate specific questions locally without waiting for text to be generated token by token. However, the reported speed depends on the device as well as the length and type of input; the response time for a single short question cannot be applied to longer text or images.

What was added since the original report

Verified updates

  1. New verified information

    Liquid AI reports a latency of 8 ms for d1-3B on NVIDIA GeForce RTX 4090.; d1-3B achieves a score of 48.57 on the public portion of Decision Index v0.2.1.; d1-omni-600M achieves a score of 15.95 on the same benchmark.

    • Liquid AI reports a latency of 8 ms for d1-3B on NVIDIA GeForce RTX 4090.
    • d1-3B achieves a score of 48.57 on the public portion of Decision Index v0.2.1.
    • d1-omni-600M achieves a score of 15.95 on the same benchmark.

Two audiences, two different impacts

What this means

01

For individuals

Developers can try decision-making based on text and images on their own devices; d1-omni-600M is available as an early research version for experiments with text and audio.

What to do Try the available demo in the System One Arcade Space on Hugging Face.
More practical updates →
02

For a business

For customer support, the demonstrated capability is to identify the topic of a request, the team responsible for handling it and its urgency in a single pass. However, the example does not demonstrate accuracy on company data or cost savings.

Processes
What to decide In a separate test, compare how d1-3B classifies anonymized customer requests with the existing team and priority assignments.
More business impacts →
d1-3B d1-omni-600M Hugging Face Liquid AI Liquid Foundation Models Nvidia Jetson

Check the original

Event sources

clearly official source · 2 publishers, 0 independent. We count feeds from the same owner only once.

2
Liquid AI (blog and model releases) primary source · first detected Open d1: Edge decision models for text, vision, and audio | Blog Hugging Face Blog primary source Multimodal open d1 decision models for the edge