Liquid AI has released two open-weight decision models: d1-3B and d1-omni-600M. They are built to answer closed-set questions (yes/no, pick one, score on a scale) in a single forward pass rather than by writing out text. Liquid says d1-3B answers one question in 16 ms on an NVIDIA Jetson AGX Thor and scores 48.57 on its Decision Index 0.2.1, ahead of every 4B and 9B model it compared and of Decider 35B-A3B at 47.11. The announcement is on the Liquid AI Hugging Face blog, and the weights are on Hugging Face.
This is a follow-up to the hosted version we covered a day earlier, Liquid AI d1 gaining image input and a claimed 19 to 200 times cost advantage over GPT-6.1 Sol. That model's weights were not stated as open. These two are.
TL;DR
| Item | d1-3B | d1-omni-600M |
|---|---|---|
| Inputs | Text and images | Text plus image, or text plus audio |
| Backbone | LFM2.5-VL-3B (decoder-only VLM) | LFM2.5-Encoder-350M (bidirectional encoder) with added vision and audio encoders |
| Status | Main release | Experimental, early research release |
| Mean score, seven public datasets | 82.9 | 78.4 |
| Speed numbers published | Yes | No |
| Where | Hugging Face | Hugging Face |

What a decision model is
A chat model produces an answer by emitting tokens one at a time. A decision model skips that. You give it a state (a support ticket, an image, a log line) and a set of named questions, and it returns structured answers directly. The question types in Liquid's example code are noul (a yes/no style check), choice (pick from labeled criteria) and score (rate on an ordered scale). Because nothing is generated, latency is dominated by one forward pass, and several questions about the same state can be answered together: Liquid reports three questions cost only about 1.3 times one question.
The category is getting crowded. Perplexity published a decisions API built on pplx-decider, OpenAI shipped a Decisions API alongside GPT-6 Luna at DevDay, and llama.cpp added decision-model support. The Liquid benchmark table compares against "Decider" models, the open-weight line these efforts have popularized. What Liquid adds is a small multimodal pair meant for devices.
How they were built
Both models sit on Liquid Foundation Models. d1-3B is trained from LFM2.5-VL-3B, the company's latest vision-language model, a decoder-only design that takes text and images. d1-omni-600M starts from LFM2.5-Encoder-350M, a bidirectional encoder, and bolts on vision and audio encoders so it can take text with either an image or audio. Liquid labels the omni model an early research release that is still being developed.
Reported benchmarks
Liquid evaluated on its Decision Index and on seven public datasets spanning reading comprehension, toxicity detection, intent classification, medical QA and cross-lingual understanding.
| Benchmark | d1-omni-600M | d1-3B | Decider 2B | Decider 4B |
|---|---|---|---|---|
| SQuAD 2.0 | 74.0 | 83.3 | 67.7 | 76.0 |
| Civil Comments | 95.8 | 93.3 | 93.6 | 92.8 |
| MASSIVE intent | 86.1 | 86.9 | 81.1 | 88.3 |
| PubMedQA | 61.3 | 68.3 | 65.7 | 63.3 |
| BoolQ | 77.7 | 86.3 | 87.3 | 89.0 |
| XNLI | 74.7 | 85.6 | 85.0 | 88.6 |
| PAWS-X | 79.5 | 76.4 | 59.5 | 69.8 |
| Mean | 78.4 | 82.9 | 77.1 | 81.1 |
The picture is more mixed than the headline. d1-3B wins the mean, but Decider 4B beats it on BoolQ, XNLI and MASSIVE intent, and the 600M omni model beats both Decider models on Civil Comments and PAWS-X. Liquid says d1-omni-600M surpasses Decider 2B with a quarter of the parameters. All numbers are Liquid's own; we have not reproduced them. Liquid also says it validated that d1-3B keeps its backbone's vision ability and that the omni model handles all three modalities, but reports no vision or audio benchmarks, since the Decision Index v0.3 has only a private vision split and audio decision benchmarks are "currently an open problem."
Speed on real hardware
Liquid worked with NVIDIA to measure d1-3B on RTX 4090 and Jetson devices, and also reports Apple and AMD results:
| Hardware | One question | Three questions | 3.4K-token state | 384px image |
|---|---|---|---|---|
| Jetson AGX Thor | 16 ms | 20 ms | 220 ms | 35 ms |
| Jetson AGX Orin 64 GB | 26 ms | 35 ms | 560 ms | 83 ms |
| Jetson Orin Nano | 50 ms | 73 ms | 1,640 ms | 202 ms |
| Apple M5 Pro | 30 ms | 41 ms | 640 ms | 62 ms |
| NVIDIA RTX 4090 | 8 ms | 21 ms | 102 ms | 17 ms |
| AMD MI325X | 9 ms | 14 ms | 44 ms | 18 ms |
Packed batches of 64 states reach roughly 38 per second on an Orin Nano, 110 on an AGX Orin, 262 on a Thor, 475 on an RTX 4090 and 1,106 on an MI325X. Speed for d1-omni-600M is not reported. The takeaway: single decisions land well under 50 ms on every device Liquid tested, but long contexts cost real time, with a 3.4K-token state taking 1.6 seconds on the smallest Jetson.
How to try it
Liquid's example needs transformers>=5.14, torch, torchvision and pillow, and the models ship their own code, so loading requires trust_remote_code=True. That flag executes code from the model repository, so review it first or pin a revision.
from transformers import AutoModel
import torch
model = AutoModel.from_pretrained(
"LiquidAI/d1-3B", trust_remote_code=True, dtype=torch.bfloat16
).to("cuda")
questions = {
"refund": {"type": "noul", "instructions": "Is the customer asking for a refund?"},
"team": {"type": "choice", "instructions": "Which team should handle this?",
"criteria": {"billing": "Charges, refunds, invoices",
"technical": "App or site faults",
"fraud": "Suspected unauthorised use"}},
}
print(model.system_one("I was charged twice this month, please refund one of them.", questions))
The same system_one call accepts an image as the whole state, and system_one_batch packs many requests without padding. Liquid also points to a System One Arcade Space on Hugging Face for demos. Mind the license: the d1-3B model page lists the license as "lfm1.0," Liquid's own license rather than a standard open-source one, so read the terms before commercial use; we have not verified what they permit.
Where this fits
Liquid's earlier edge releases set the stage: the LFM2.5 230M edge agent model, the LFM2.5 2.6B on-device agents and the Pipette on-device benchmark. The pattern is a small model for a narrow, high-volume job. Typical fits for a decision model:
- Routing and triage. Classify tickets, messages or events and send them to the right queue, with confidence scores for thresholds.
- Moderation. Toxicity and policy checks at the edge, with no round trip to a cloud API.
- Robotics and cameras. Ask a few yes/no questions of each frame on a Jetson without running a full VLM.
- Agent control loops. Cheap, fast decisions inside a loop, leaving the large model for planning.
Where it does not fit: open-ended generation, explanation, or any task whose answer is not a fixed set of choices.
Why single-pass decisions matter on a device
Edge hardware punishes generation. A generative model answering "which team should handle this?" has to produce a sentence, and every token costs a full pass through the network plus memory traffic that small boards like the Orin Nano have little room for. A decision model collapses that to one pass whose cost is mostly the input length. That is why Liquid's table shows the cost of a single question barely rising for three questions (16 ms to 20 ms on Thor) while a 3.4K-token state costs far more: the state is the expensive part, not the answer.
There is a second benefit that matters for product teams: structure. Because the output is a choice or a score rather than prose, there is nothing to parse and no chance of a malformed answer. A refund check either fires or it does not, and a score can be thresholded. That makes the model easy to slot into ordinary software, where a human-readable explanation is usually not wanted.
Timeline: Liquid's run of releases
For readers following Liquid AI, the last few months have been busy. The company shipped small on-device agent models in June and August, published its Pipette on-device benchmark, and this week moved from a hosted d1 with image input to open weights for the family. The consistent bet is that many production tasks are narrow and high-volume, and that a few billion parameters tuned for that task beat a frontier model on cost and latency. Open weights extend that bet to teams that cannot send data to a hosted API, such as factories, vehicles and clinics.
How to evaluate it before you adopt it
- Build a small labeled set from your own data. Public benchmarks such as BoolQ or Civil Comments say little about your ticket taxonomy or camera angles. Fifty to two hundred labeled examples is enough for a first signal.
- Compare against a baseline you already have. That might be a fine-tuned small classifier, a rules engine or a prompted LLM. Look at accuracy, latency and cost together.
- Check calibration. If you plan to route by confidence, test whether a 0.9 really means 90% on your data.
- Measure on the target device. Liquid's latencies are for its test setups; quantization, power mode and input size on your board will differ.
- Read the license and the remote code. The
trust_remote_coderequirement and the custom license both deserve a look from whoever signs off on dependencies. - Treat the omni model as a preview. Liquid says it is still being developed and published no speed or modality benchmarks for it.
What to watch
Independent reproductions of the Decision Index numbers, real-world accuracy on your own labels, how the omni model matures past its early research stage, and how the license is read in practice. Download counts on the model page were tiny at publication, so community feedback is still to come.
Related reading
- Liquid AI d1 vision: the hosted model and its cost claims
- Perplexity pplx-decider Decisions API
- llama.cpp decision model support
- Liquid AI LFM2.5 2.6B for on-device agents
- Edge0 35B on-device model on Apple Silicon
Figures are vendor-reported as of October 7, 2026 and may be revised.
