2026-10-07

Liquid AI Releases Open d1, Two Open-Weight Decision Models: d1-3B Scores 48.6 on the Decision Index and Answers in 8 ms on an RTX 4090

AIModelsOpen Source🌍 North America

Liquid AI released Open d1 on October 7, 2026: two open-weight models from its d1 decision family, with weights on Hugging Face. d1-3B takes text and vision, and d1-omni-600M takes text plus either an image or audio. Decision models return an answer in one forward pass rather than generating tokens. On the public split of Decision Index v0.2.1, d1-3B scores 48.57, ahead of every model under 10B and on par with Decider 35B-A3B, a decision model 12x its size. The experimental d1-omni-600M scores 15.95. Liquid's blog post positions the pair for everything from NVIDIA DGX in the data center to Jetson at the edge.

Decision models, not generators

Unlike Liquid's generative LFMs, d1 models output a decision: a yes/no answer with a probability, a choice among options with a confidence and per-criterion probabilities, or a score with a confidence. In the diagrams, a state plus an instruction template goes in, and a head reads the decision off the final layer.

d1-3B starts from LFM2.5-VL-3B. Liquid averaged the weights of LFM2.5-2.6B with that model's text backbone for a stronger base, fine-tuned checkpoints with different seeds and data mixtures, then merged them again. Training on long inputs, shuffling answer options and fixing shortcuts in the data mattered more than more advanced techniques.

Architecture of d1-3B: a SigLIP2 vision encoder and tokenizer feed an LFM2.5-2.6B causal backbone, whose LM head produces yes/no, choice and score outputs. From Liquid AI's blog post.

d1-omni-600M is built on LFM2.5-Encoder-350M, a bidirectional encoder, with a SigLIP2 vision encoder or a FastConformer audio encoder attached. It was fine-tuned on decision tasks first, then audio and vision were added in stages: the audio encoder was tuned against a frozen text backbone, and the vision encoder from LFM2.5-VL-450M was added with an adapter and LoRA updates active only when images are present. A full fine-tune, a LoRA merge and weight averaging closed the process.

Architecture of d1-omni-600M: image or audio encoders feed a 16-layer bidirectional LFM2.5-Encoder-350M, followed by a decision head. From Liquid AI's blog post.

Benchmarks

On the Decision Index chart, d1-3B sits at roughly 48.5 at 3B parameters, on the size-versus-score frontier, and d1-omni-600M at roughly 16 at 0.6B. Fine-tuned and inference-technique entrants up to 32B land between roughly 40 and 57, and two hosted APIs sit higher, d1 at roughly 60 and Jev 1.13 at roughly 57 to 58.

Decision Index 0.2.1 against number of parameters on a log axis, with d1-3B and d1-omni-600M highlighted and dashed lines for two hosted APIs. From Liquid AI's thread.

On seven public text benchmarks, d1-3B has the highest mean, 82.9 against 81.1 for Decider 4B. d1-omni-600M's 78.4 beats Decider 2B's 77.1 at 0.6B parameters, under a third of its size.

Benchmarkd1-omni-600Md1-3BDecider 2BDecider 4B
SQuAD 2.074.085.367.776.0
Civil Comments95.893.093.692.8
MASSIVE intent86.187.381.188.3
PubMedQA61.366.065.763.3
BoolQ77.786.787.389.0
XNLI74.785.085.088.6
PAWS-X79.576.959.569.8
Mean78.482.977.181.1

The means recompute exactly, but the mean hides the row picture. d1-3B tops only two rows, SQuAD 2.0 and PubMedQA; Decider 4B wins MASSIVE, BoolQ (89.0 against 86.7) and XNLI (88.6 against 85.0), and d1-3B only ties Decider 2B on XNLI at 85.0. The small omni model takes the other two, Civil Comments (toxicity) and PAWS-X (paraphrase).

Decision Index v0.3 has a private vision split that is not reported; Liquid validated that d1-3B keeps LFM2.5-VL-3B's vision capabilities on standard benchmarks and that d1-omni-600M handles all three modalities in its playground demos. Dedicated audio decision benchmarks do not exist yet, and Liquid hopes the community builds them.

Fast on edge hardware

Because nothing is generated, latency is end-to-end. Liquid's d1-3B figures, one request at a time (edge devices measured with NVIDIA):

Device1 question3 questions3.4K-token state384px image64 states packed
RTX 40908 ms21 ms102 ms17 ms475/s
AMD MI325X9 ms14 ms44 ms18 ms1,106/s
Jetson AGX Thor16 ms20 ms220 ms35 ms262/s
Jetson AGX Orin 64GB26 ms35 ms560 ms83 ms110/s
Apple M5 Pro30 ms41 ms640 ms62 ms78/s
Jetson Orin Nano50 ms73 ms1,640 ms202 ms38/s

A single question stays at or under 50 ms on every device, and under 10 ms on both GPUs. Three questions cost about 1.3x one question on the edge devices (Thor goes from 16 to 20 ms, a 1.25x ratio); the RTX 4090 pays about 2.6x. llama.cpp support with NVFP4 shipped on day one across Apple, AMD, Qualcomm and NVIDIA. Numbers are for d1-3B only, as d1-omni-600M is an early research release.

What you can build

Liquid published ten demos in a Hugging Face space, with no setup, that run d1-3B in a loop over live camera input, one pass per frame. They range from gesture-controlled games to live content moderation. With NVIDIA, Liquid also showed d1-3B navigating an environment in Isaac Sim, served from a Jetson in a hardware-in-the-loop setup.

Context

d1-3B builds on LFM2.5-VL-3B, Liquid's vision-language model, and on the weights of LFM2.5-2.6B, the small model Liquid pitched for local agents. NVIDIA's own small specialist worker is covered in Nemotron 3.5 Lightning, and EmbeddingGemma 2 is another multimodal model sized for on-device use.

Read next