NEWVectors or files. Pick a path.Start →
    Models/MCG-NJU/OneStreamer-4B
    Apache 2.0

    OneStreamer-4B

    by MCG-NJU

    A 4B streaming video model that keeps time-stamped memory and decides when to speak, 72.1 on OVOBench

    9likes
    4.4B (Qwen3-VL-4B base)params
    Identifiers
    Model ID
    MCG-NJU/OneStreamer-4B
    Feature URI

    Deploy OneStreamer-4B

    Single-tenant

    Mixpeek has no managed extractor for this model. On a single-tenant deployment you upload the weights and a custom plugin serves them next to the rest of your pipeline.

    Overview

    OneStreamer-4B is a streaming video-language model from Nanjing University's MCG lab, built on Qwen3-VL-4B-Instruct and released under Apache 2.0 in September 2026. It reads video as it arrives through a recent visual window, keeps a time-aligned caption memory of what it has seen, and decides on its own when it has enough evidence to answer.\n\nWhile watching it emits a silence token to keep observing, a standby token when it needs more evidence, or a response token followed by an answer. The card reports gains over the Qwen3-VL-4B base on eight streaming-video benchmarks, including 72.1 against 58.8 on OVOBench and 48.7 against 34.3 on ProactiveVideoQA.

    Architecture

    Qwen3-VL-4B-Instruct (about 4.4B parameters) trained for streaming interaction. Video is processed incrementally through a recent visual window, with PHCM retaining a time-aligned caption memory for later questions. Proactive responses are controlled by Silence, Standby and Response tokens. Training data is the OneStreamer-1M dataset.

    Mixpeek SDK Integration

    # OneStreamer runs on your side with the Inference/ code in its GitHub repository.
    # Store each time-stamped caption it writes as a text object, so a Mixpeek text
    # collection can search the stream later.
    import requests
    
    for cap in captions:  # e.g. {"start": 1312.0, "end": 1318.5, "text": "A forklift enters aisle 4"}
        requests.post(
            "https://api.mixpeek.com/v1/buckets/bkt_your_bucket/objects",
            headers={"Authorization": "Bearer API_KEY", "X-Namespace": "ns_your_namespace"},
            json={"blobs": [{"property": "caption", "type": "text", "data": cap["text"]}],
                  "metadata": {"stream": "dock-cam-2", "start": cap["start"], "end": cap["end"]}},
        )

    Capabilities

    • Streaming video input through a recent visual window
    • Time-aligned caption memory for questions about earlier moments
    • Proactive replies: decides when it has seen enough to answer
    • English and Chinese
    • Apache 2.0 license

    Use Cases on Mixpeek

    Live camera or broadcast monitoring that answers when an event happens
    Assistants that answer questions about something shown minutes earlier
    Producing time-stamped captions of a stream for later search

    Benchmarks

    DatasetMetricScoreSource
    OVOBenchOverall72.1Model card: MCG-NJU/OneStreamer-4B (self-reported; Qwen3-VL-4B base 58.8)
    StreamingBenchReal-Time86.9Model card (self-reported; base 81.8)
    ODVBenchOverall71.3Model card (self-reported; base 57.6)
    ProactiveVideoQAAverage48.7Model card (self-reported; base 34.3)
    OVO-TimingAverage F141.6Model card (self-reported; base 29.4)

    Performance

    Input SizeStreaming video frames through a recent window, with text questions in English or Chinese
    Embedding Dimn/a (outputs text)
    GPU LatencyInput dependent
    GPU ThroughputBatch dependent
    GPU MemoryModel dependent

    Scores are the card's, each benchmark on its own metric. The OmniMMI result uses speech recognition for some subtasks. We have not measured it.

    Frequently Asked Questions

    What is OneStreamer-4B?

    A streaming video-language model built on Qwen3-VL-4B-Instruct. It watches video as it arrives, keeps a time-aligned memory of what it saw, and decides when to answer.

    How is a streaming video model different from a normal video model?

    A normal video model reads a finished clip and then answers. A streaming model reads frames as they arrive, has to remember earlier moments without re-reading them, and here also chooses when to speak, which matters for live feeds.

    How well does OneStreamer-4B score?

    The card reports 72.1 on OVOBench, 86.9 on StreamingBench real-time and 48.7 on ProactiveVideoQA, each well above the Qwen3-VL-4B base model.

    How do I search what OneStreamer saw with Mixpeek?

    Store each time-stamped caption it writes as a text object with the stream name and times in metadata, then search them with a text retriever, as in the example on this page.

    Specification

    OrganizationMCG-NJU
    Retriever-
    Parameters4.4B (Qwen3-VL-4B base)
    LicenseApache 2.0
    Downloads/moN/A
    Likes9

    Research Paper

    OneStreamer: Unifying Perception, Memory, and Proactive Response in Streaming Video Interaction

    arxiv.org

    Build a pipeline with OneStreamer-4B

    Add this model to a processing pipeline alongside other extractors. Combine with retrieval stages for end-to-end search.

    Run it on your own data, free