TimeLens2-8B
by MCG-NJU
An 8B video model that answers a text query with the exact seconds where it happens, at 48.0 average mIoU on seven grounding benchmarks
MCG-NJU/TimeLens2-8BDeploy TimeLens2-8B
Single-tenantMixpeek has no managed extractor for this model. On a single-tenant deployment you upload the weights and a custom plugin serves them next to the rest of your pipeline.
Overview
TimeLens2-8B is a video temporal grounding model from Nanjing University's MCG lab, fine-tuned from Qwen3-VL-8B-Instruct. Give it a video and a description such as "a man opens the refrigerator" and it returns every time span where that happens, as start and end seconds.
The card reports 48.0 average mIoU across seven temporal grounding benchmarks covering short, long and egocentric video, and calls that a new state of the art on the suite. It was released in July 2026 with 2B and 4B siblings, under Apache-2.0.
Architecture
A Qwen3-VL-8B multimodal LLM fine-tuned for temporal grounding. Frames are sampled from the video (2 fps in the card's example), encoded with the Qwen3-VL vision encoder, and the language model is prompted to return a JSON array of [start, end] pairs in seconds for the query.
Mixpeek SDK Integration
# Index the library first: a Mixpeek video collection cuts each file into scenes
# and makes them searchable. TimeLens2 then refines a found scene to exact seconds.
import requests
requests.post(
"https://api.mixpeek.com/v1/buckets/bkt_your_bucket/objects",
headers={"Authorization": "Bearer API_KEY", "X-Namespace": "ns_your_namespace"},
json={"blobs": [{"property": "video", "type": "video", "data": "s3://footage/warehouse-cam-07.mp4"}]},
)Capabilities
- Returns every time span in a video that matches a text description, as [start, end] seconds
- Short, long and first-person (egocentric) video
- Runs with Hugging Face transformers on one GPU
- Apache-2.0, so commercial use is allowed
Use Cases on Mixpeek
Benchmarks
| Dataset | Metric | Score | Source |
|---|---|---|---|
| Seven temporal grounding benchmarks (average) | mIoU | 48.0 | Model card: MCG-NJU/TimeLens2-8B (self-reported, reported as state of the art on that suite) |
Performance
A generative model: it reads the whole clip for every query, so it is too slow to search a library on its own. Use it on the few clips a retriever returns. The card does not break the 48.0 average out by benchmark in text. We have not measured it.
Common Pipeline Companions
Frequently Asked Questions
What does TimeLens2 do?
It finds when something happens inside a video. Given a clip and a text description, it returns the time spans, in seconds, where the description is true.
Can TimeLens2 search a whole video library?
Not on its own. It reads the whole clip for each query, so it suits the handful of clips a search returns. Index the library with an embedding-based search first, then run TimeLens2 on the top results.
Can I use TimeLens2 commercially?
Yes. It is released under Apache-2.0.
How do I use TimeLens2 with Mixpeek?
Index the footage in a Mixpeek video collection and search it with a retriever, which returns scenes with timestamps. Run TimeLens2 on the top scene with the same query when you need tighter start and end times than the scene boundaries.
Specification
Research Paper
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs
arxiv.orgBuild a pipeline with TimeLens2-8B
Add this model to a processing pipeline alongside other extractors. Combine with retrieval stages for end-to-end search.
Run it on your own data, free