OneStreamer-4B
by MCG-NJU
A 4B streaming video model that keeps time-stamped memory and decides when to speak, 72.1 on OVOBench
MCG-NJU/OneStreamer-4BDeploy OneStreamer-4B
Single-tenantMixpeek has no managed extractor for this model. On a single-tenant deployment you upload the weights and a custom plugin serves them next to the rest of your pipeline.
Overview
OneStreamer-4B is a streaming video-language model from Nanjing University's MCG lab, built on Qwen3-VL-4B-Instruct and released under Apache 2.0 in September 2026. It reads video as it arrives through a recent visual window, keeps a time-aligned caption memory of what it has seen, and decides on its own when it has enough evidence to answer.\n\nWhile watching it emits a silence token to keep observing, a standby token when it needs more evidence, or a response token followed by an answer. The card reports gains over the Qwen3-VL-4B base on eight streaming-video benchmarks, including 72.1 against 58.8 on OVOBench and 48.7 against 34.3 on ProactiveVideoQA.
Architecture
Qwen3-VL-4B-Instruct (about 4.4B parameters) trained for streaming interaction. Video is processed incrementally through a recent visual window, with PHCM retaining a time-aligned caption memory for later questions. Proactive responses are controlled by Silence, Standby and Response tokens. Training data is the OneStreamer-1M dataset.
Mixpeek SDK Integration
# OneStreamer runs on your side with the Inference/ code in its GitHub repository.
# Store each time-stamped caption it writes as a text object, so a Mixpeek text
# collection can search the stream later.
import requests
for cap in captions: # e.g. {"start": 1312.0, "end": 1318.5, "text": "A forklift enters aisle 4"}
requests.post(
"https://api.mixpeek.com/v1/buckets/bkt_your_bucket/objects",
headers={"Authorization": "Bearer API_KEY", "X-Namespace": "ns_your_namespace"},
json={"blobs": [{"property": "caption", "type": "text", "data": cap["text"]}],
"metadata": {"stream": "dock-cam-2", "start": cap["start"], "end": cap["end"]}},
)Capabilities
- Streaming video input through a recent visual window
- Time-aligned caption memory for questions about earlier moments
- Proactive replies: decides when it has seen enough to answer
- English and Chinese
- Apache 2.0 license
Use Cases on Mixpeek
Benchmarks
| Dataset | Metric | Score | Source |
|---|---|---|---|
| OVOBench | Overall | 72.1 | Model card: MCG-NJU/OneStreamer-4B (self-reported; Qwen3-VL-4B base 58.8) |
| StreamingBench | Real-Time | 86.9 | Model card (self-reported; base 81.8) |
| ODVBench | Overall | 71.3 | Model card (self-reported; base 57.6) |
| ProactiveVideoQA | Average | 48.7 | Model card (self-reported; base 34.3) |
| OVO-Timing | Average F1 | 41.6 | Model card (self-reported; base 29.4) |
Performance
Scores are the card's, each benchmark on its own metric. The OmniMMI result uses speech recognition for some subtasks. We have not measured it.
Common Pipeline Companions
Frequently Asked Questions
What is OneStreamer-4B?
A streaming video-language model built on Qwen3-VL-4B-Instruct. It watches video as it arrives, keeps a time-aligned memory of what it saw, and decides when to answer.
How is a streaming video model different from a normal video model?
A normal video model reads a finished clip and then answers. A streaming model reads frames as they arrive, has to remember earlier moments without re-reading them, and here also chooses when to speak, which matters for live feeds.
How well does OneStreamer-4B score?
The card reports 72.1 on OVOBench, 86.9 on StreamingBench real-time and 48.7 on ProactiveVideoQA, each well above the Qwen3-VL-4B base model.
How do I search what OneStreamer saw with Mixpeek?
Store each time-stamped caption it writes as a text object with the stream name and times in metadata, then search them with a text retriever, as in the example on this page.
Specification
Research Paper
OneStreamer: Unifying Perception, Memory, and Proactive Response in Streaming Video Interaction
arxiv.orgBuild a pipeline with OneStreamer-4B
Add this model to a processing pipeline alongside other extractors. Combine with retrieval stages for end-to-end search.
Run it on your own data, free