NEWVectors or files. Pick a path.Start →
    Models/LiquidAI/d1-omni-600M
    LFM Open License v1.0 (commercial use under $10M annual revenue)

    d1-omni-600M

    by LiquidAI

    A 587M open decision model that answers questions about text, images or a voice clip in one pass, sized for edge devices

    Identifiers
    Model ID
    LiquidAI/d1-omni-600M
    Feature URI

    Deploy d1-omni-600M

    Single-tenant

    Mixpeek has no managed extractor for this model. On a single-tenant deployment you upload the weights and a custom plugin serves them next to the rest of your pipeline.

    Overview

    d1-omni-600M is the small member of Liquid AI's open d1 decision family, released in October 2026. It reads text together with images, or text together with up to 30 seconds of speech, and answers named questions (yes or no, pick one, or a rating) as probabilities, with no generated text.

    At 587M parameters it targets phones, wearables and other edge devices. It is strongest on narrow checks: Liquid reports 95.8 on Civil Comments moderation, above d1-3B, but 15.95 on the broad Decision Index, where d1-3B scores 48.57. Liquid calls it an early research release.

    Architecture

    A 381M shared trunk and decision head built on LFM2.5-Encoder-350M, a 94M SigLIP2 vision encoder from LFM2.5-VL-450M and a 112M, 17-layer FastConformer audio encoder. Every modality runs through the same trunk weights, and each answer is read from the distribution over the options.

    Mixpeek SDK Integration

    # Index the recordings or images in Mixpeek; run d1-omni on the device that captures them.
    import requests
    
    requests.post(
        "https://api.mixpeek.com/v1/buckets/bkt_your_bucket/objects",
        headers={"Authorization": "Bearer API_KEY", "X-Namespace": "ns_your_namespace"},
        json={"blobs": [{"property": "video", "type": "video", "data": "s3://store-cams/aisle-4.mp4"}]},
    )

    Capabilities

    • Yes/no, pick-one and rating questions answered as probabilities, with zero output tokens
    • Text with images, or text with up to 30 seconds of speech, in one pass
    • 587M parameters, sized for edge devices
    • 16,384-token context across text, image and audio positions

    Use Cases on Mixpeek

    Routing voice commands on a device without transcribing them first
    Moderation checks on text and images at the edge
    Intent and topic classification where a 3B model will not fit

    Benchmarks

    DatasetMetricScoreSource
    Civil CommentsScore95.8Model card: LiquidAI/d1-omni-600M (Liquid's internal evaluation; d1-3B 93.0)
    7 public benchmarks as decisionsMean78.4Model card (d1-3B 82.9 on the same set)
    Fast Decisions (dev)Score76.9Model card (self-reported)
    Decision Index 0.2.1Score15.95Model card (d1-3B 48.57)

    Performance

    Input SizeText with images, or text with up to 30 s of speech; 16,384-token context
    Embedding Dimn/a (returns a probability per option)
    GPU LatencyInput dependent
    GPU ThroughputBatch dependent
    GPU MemoryModel dependent

    Liquid publishes no speed figures for this early research release. Audio was trained on English requests to an assistant, and clips are cut at 30 seconds. With images, the state and question text is cut to 896 tokens. Run it in float32: Liquid found bfloat16 changed the top answer on 0.8% of text and 1.7% of audio rows. We have not measured it.

    Frequently Asked Questions

    What is d1-omni-600M for?

    Fast yes/no, choice and rating decisions on small devices, over text with images or text with a short voice clip: routing voice commands, moderation and intent classification.

    Can d1-omni-600M understand speech?

    It reads up to 30 seconds of audio together with text and answers questions about it, such as what the speaker wants. Liquid trained the audio side on English requests to an assistant, so test other speech before relying on it.

    How does d1-omni-600M compare with d1-3B?

    d1-3B is stronger on broad decisions (48.57 against 15.95 on the Decision Index) and on most benchmarks, but it handles text and images only. d1-omni-600M is about a fifth of the size and also takes audio.

    Can I use d1-omni-600M commercially?

    Under the LFM Open License v1.0, commercial use is allowed for organisations with less than $10M in annual revenue; above that you need an agreement with Liquid AI.

    Specification

    OrganizationLiquidAI
    Retriever-
    Parameters587M
    LicenseLFM Open License v1.0 (commercial use under $10M annual revenue)
    Downloads/moN/A
    Likes50

    Research Paper

    Open d1: Edge decision models for text, vision, and audio (Liquid AI)

    arxiv.org

    Build a pipeline with d1-omni-600M

    Add this model to a processing pipeline alongside other extractors. Combine with retrieval stages for end-to-end search.

    Run it on your own data, free