Model News

Liquid AI Releases Multimodal Open D1 Decision Models for Edge

Liquid AI released d1-3B and d1-omni-600M, open multimodal decision models designed for edge devices, supporting text, images, and audio inputs.

Headline card: Liquid AI Releases Multimodal Open D1 Decision Models for Edge
On this page
  1. What changed
  2. Why it matters
  3. What to test
  4. The conclusion

What changed

Liquid AI released two open decision models on October 7, 2026, extending the decision model category into multimodal territory. The models are d1-3B and d1-omni-600M (experimental), both available as open weights on Hugging Face.

Decision models differ from generative models. Instead of producing tokens sequentially, they answer in a single forward pass, designed to handle structured decision tasks like classification, scoring, or choice selection. This architecture makes them suitable for edge deployment where latency matters.

The d1-3B model accepts text and images as inputs. The d1-omni-600M accepts either text and images or text and audio, though it remains in early research release. Both are built on Liquid Foundation Models (LFMs). The d1-3B uses LFM2.5-VL-3B, a decoder-only vision-language model. The d1-omni-600M uses LFM2.5-Encoder-350M, a bidirectional encoder, with added vision and audio encoders for multimodal support.

Why it matters

This announcement expands what decision models can handle. Until now, the category focused primarily on text-based tasks. Adding vision and audio support opens decision models to new use cases on edge devices where computational resources are constrained.

Speed matters for edge deployment. According to the announcement, d1-3B answers a single question in under 50 milliseconds on every measured device, from consumer hardware like Apple M5 Pro to embedded systems like Jetson Orin Nano. Batch processing three questions takes only 1.3x the time of one on the AGX Thor, suggesting efficiency gains at scale.

The performance claims are noteworthy. Liquid AI says d1-3B achieves a mean score of 82.9 across seven public benchmarks, higher than Decider 4B (81.1). Meanwhile, d1-omni-600M scores 78.4 with a quarter of Decider 2B’s parameters while surpassing its performance of 77.1. These results span reading comprehension, toxicity detection, intent classification, medical QA, and cross-lingual understanding.

Who benefits? Organizations building on-device applications need structured decision outputs rather than free-form text. Chatbot routing, content moderation, customer service ticket triage, and medical question answering are examples where decision models fit. The open weights mean researchers and companies can run these models locally without cloud API dependencies, keeping data on-device and controlling costs.

The source confirms that Liquid AI validated d1-3B retains vision capabilities of its LFM2.5-VL-3B backbone on standard vision benchmarks, and that d1-omni-600M handles all three modalities. However, the announcement does not report specific vision or audio benchmark scores, noting that audio decision benchmarks remain an open problem.

What to test

Before adopting, several vendor claims require verification in your own environment.

Test d1-3B’s stated response times on your specific hardware. The announcement reports benchmarks across multiple devices, but real-world latency depends on batch size, input complexity, and integration overhead. Run inference on your target device with production-scale inputs. The source shows processing a 3.4K-token state takes 220 ms on a Jetson AGX Thor, much slower than single-question inference at 16 ms. If your use case involves long context windows, measure latency with realistic state sizes.

The models were tested on seven datasets spanning reading comprehension, toxicity detection, intent classification, medical QA, and cross-lingual understanding. These may not reflect your task distribution. The announcement does not report vision or audio decision benchmarks, acknowledging that audio decision benchmarks remain an open problem. Test on held-out data matching your domain.

While d1-3B’s vision capabilities were validated against standard vision benchmarks, the announcement does not disclose those specific results. For d1-omni-600M, no audio decision benchmarks are reported. Evaluate vision and audio performance directly before production deployment.

d1-omni-600M is explicitly flagged as experimental and undergoing further development. Production stability and future API changes are unknowns. Use in experimental systems only until Liquid AI signals production readiness.

The models require trust_remote_code=True, meaning custom inference logic runs outside standard transformers. Audit this code before deployment in secure environments.

The conclusion

Liquid AI’s release of multimodal open d1 decision models extends decision models into new territory, adding vision and audio support for edge deployment. The speed metrics and benchmark results are encouraging for structured decision tasks, but the field remains young. Audio decision benchmarking is an acknowledged open problem, and d1-omni-600M’s experimental status warrants caution.

The practical value is clear for use cases requiring fast, structured outputs with multimodal inputs on constrained hardware. The open weights lower barriers to adoption compared to API-dependent alternatives. However, vision and audio benchmarks are not disclosed, and generalization to novel tasks remains untested at scale. Early adopters should pilot these models on non-critical workloads first, testing latency claims and decision quality against their specific data and hardware.

Watch for future releases that add audio benchmarks, move d1-omni-600M to stable status, and expand the decision model family to new modalities. The decision model category is still defining itself, and this announcement signals progress toward practical multimodal on-device AI.

AI Tool Herald may earn a commission from some links on this site. It never changes what we report or recommend. Affiliate disclosure