Liquid AI releases weights for d1-3B and d1-omni-600M for local decision-making
Liquid AI has released the open weights of d1-3B and d1-omni-600M. The models return decisions in a single pass without generating tokens. They support text with images, and the experimental d1-omni-600M also supports text with audio.
Liquid AI has made the open weights of the decision models d1-3B and d1-omni-600M available on Hugging Face for local use. d1-3B accepts text and images, while d1-omni-600M accepts a combination of text and an image or text and audio. According to the company, the models return an answer in a single pass through the network without generating tokens. d1-omni-600M is an early research version that is still under development.
For d1-3B, the company reports a response time for a single question of 8 ms on NVIDIA GeForce RTX 4090, 16 ms on NVIDIA Jetson AGX Thor, 26 ms on NVIDIA Jetson AGX Orin 64 GB and 50 ms on NVIDIA Jetson Orin Nano. Each measurement used a single request. Longer inputs have significantly higher latency: on NVIDIA Jetson Orin Nano, an input of 3 400 tokens took 1 640 ms, and processing a 384 px image took 202 ms. The company has not published speed figures for d1-omni-600M.
According to the company, d1-3B achieved a score of 48.57 on the public portion of Decision Index v0.2.1, and d1-omni-600M scored 15.95. Both articles report averages across seven text benchmarks of 82.9 and 78.4, respectively, but some individual results differ. For example, for d1-3B on SQuAD 2.0, the Liquid AI blog reports 85.3, while the article on Hugging Face reports 83.3. The published results do not include a separate evaluation of decision-making based on images and audio.
The models have had llama.cpp support since release. The article on Hugging Face also includes an example using the transformers library, version 5.14 or later, which simultaneously identifies a refund request, the appropriate team and the urgency within a single customer request. Demos are available to try in the System One Arcade Space on Hugging Face.
Why it matters
Developers of applications that respond to camera images can evaluate specific questions locally without waiting for text to be generated token by token. However, the reported speed depends on the device as well as the length and type of input; the response time for a single short question cannot be applied to longer text or images.
What was added since the original report
Verified updates
-
Liquid AI reports a latency of 8 ms for d1-3B on NVIDIA GeForce RTX 4090.; d1-3B achieves a score of 48.57 on the public portion of Decision Index v0.2.1.; d1-omni-600M achieves a score of 15.95 on the same benchmark.
- Liquid AI reports a latency of 8 ms for d1-3B on NVIDIA GeForce RTX 4090.
- d1-3B achieves a score of 48.57 on the public portion of Decision Index v0.2.1.
- d1-omni-600M achieves a score of 15.95 on the same benchmark.
Two audiences, two different impacts
What this means
For individuals
Developers can try decision-making based on text and images on their own devices; d1-omni-600M is available as an early research version for experiments with text and audio.
For a business
For customer support, the demonstrated capability is to identify the topic of a request, the team responsible for handling it and its urgency in a single pass. However, the example does not demonstrate accuracy on company data or cost savings.
ProcessesCheck the original
Event sources
clearly official source · 2 publishers, 0 independent. We count feeds from the same owner only once.