d1-omni-600M
by LiquidAI
A 587M open decision model that answers questions about text, images or a voice clip in one pass, sized for edge devices
LiquidAI/d1-omni-600MDeploy d1-omni-600M
Single-tenantMixpeek has no managed extractor for this model. On a single-tenant deployment you upload the weights and a custom plugin serves them next to the rest of your pipeline.
Overview
d1-omni-600M is the small member of Liquid AI's open d1 decision family, released in October 2026. It reads text together with images, or text together with up to 30 seconds of speech, and answers named questions (yes or no, pick one, or a rating) as probabilities, with no generated text.
At 587M parameters it targets phones, wearables and other edge devices. It is strongest on narrow checks: Liquid reports 95.8 on Civil Comments moderation, above d1-3B, but 15.95 on the broad Decision Index, where d1-3B scores 48.57. Liquid calls it an early research release.
Architecture
A 381M shared trunk and decision head built on LFM2.5-Encoder-350M, a 94M SigLIP2 vision encoder from LFM2.5-VL-450M and a 112M, 17-layer FastConformer audio encoder. Every modality runs through the same trunk weights, and each answer is read from the distribution over the options.
Mixpeek SDK Integration
# Index the recordings or images in Mixpeek; run d1-omni on the device that captures them.
import requests
requests.post(
"https://api.mixpeek.com/v1/buckets/bkt_your_bucket/objects",
headers={"Authorization": "Bearer API_KEY", "X-Namespace": "ns_your_namespace"},
json={"blobs": [{"property": "video", "type": "video", "data": "s3://store-cams/aisle-4.mp4"}]},
)Capabilities
- Yes/no, pick-one and rating questions answered as probabilities, with zero output tokens
- Text with images, or text with up to 30 seconds of speech, in one pass
- 587M parameters, sized for edge devices
- 16,384-token context across text, image and audio positions
Use Cases on Mixpeek
Benchmarks
| Dataset | Metric | Score | Source |
|---|---|---|---|
| Civil Comments | Score | 95.8 | Model card: LiquidAI/d1-omni-600M (Liquid's internal evaluation; d1-3B 93.0) |
| 7 public benchmarks as decisions | Mean | 78.4 | Model card (d1-3B 82.9 on the same set) |
| Fast Decisions (dev) | Score | 76.9 | Model card (self-reported) |
| Decision Index 0.2.1 | Score | 15.95 | Model card (d1-3B 48.57) |
Performance
Liquid publishes no speed figures for this early research release. Audio was trained on English requests to an assistant, and clips are cut at 30 seconds. With images, the state and question text is cut to 896 tokens. Run it in float32: Liquid found bfloat16 changed the top answer on 0.8% of text and 1.7% of audio rows. We have not measured it.
Common Pipeline Companions
Frequently Asked Questions
What is d1-omni-600M for?
Fast yes/no, choice and rating decisions on small devices, over text with images or text with a short voice clip: routing voice commands, moderation and intent classification.
Can d1-omni-600M understand speech?
It reads up to 30 seconds of audio together with text and answers questions about it, such as what the speaker wants. Liquid trained the audio side on English requests to an assistant, so test other speech before relying on it.
How does d1-omni-600M compare with d1-3B?
d1-3B is stronger on broad decisions (48.57 against 15.95 on the Decision Index) and on most benchmarks, but it handles text and images only. d1-omni-600M is about a fifth of the size and also takes audio.
Can I use d1-omni-600M commercially?
Under the LFM Open License v1.0, commercial use is allowed for organisations with less than $10M in annual revenue; above that you need an agreement with Liquid AI.
Specification
Research Paper
Open d1: Edge decision models for text, vision, and audio (Liquid AI)
arxiv.orgBuild a pipeline with d1-omni-600M
Add this model to a processing pipeline alongside other extractors. Combine with retrieval stages for end-to-end search.
Run it on your own data, free