NEWVectors or files. Pick a path.Start →
    Models/MCG-NJU/TimeLens2-8B
    Apache-2.0

    TimeLens2-8B

    by MCG-NJU

    An 8B video model that answers a text query with the exact seconds where it happens, at 48.0 average mIoU on seven grounding benchmarks

    Identifiers
    Model ID
    MCG-NJU/TimeLens2-8B
    Feature URI

    Deploy TimeLens2-8B

    Single-tenant

    Mixpeek has no managed extractor for this model. On a single-tenant deployment you upload the weights and a custom plugin serves them next to the rest of your pipeline.

    Overview

    TimeLens2-8B is a video temporal grounding model from Nanjing University's MCG lab, fine-tuned from Qwen3-VL-8B-Instruct. Give it a video and a description such as "a man opens the refrigerator" and it returns every time span where that happens, as start and end seconds.

    The card reports 48.0 average mIoU across seven temporal grounding benchmarks covering short, long and egocentric video, and calls that a new state of the art on the suite. It was released in July 2026 with 2B and 4B siblings, under Apache-2.0.

    Architecture

    A Qwen3-VL-8B multimodal LLM fine-tuned for temporal grounding. Frames are sampled from the video (2 fps in the card's example), encoded with the Qwen3-VL vision encoder, and the language model is prompted to return a JSON array of [start, end] pairs in seconds for the query.

    Mixpeek SDK Integration

    # Index the library first: a Mixpeek video collection cuts each file into scenes
    # and makes them searchable. TimeLens2 then refines a found scene to exact seconds.
    import requests
    
    requests.post(
        "https://api.mixpeek.com/v1/buckets/bkt_your_bucket/objects",
        headers={"Authorization": "Bearer API_KEY", "X-Namespace": "ns_your_namespace"},
        json={"blobs": [{"property": "video", "type": "video", "data": "s3://footage/warehouse-cam-07.mp4"}]},
    )

    Capabilities

    • Returns every time span in a video that matches a text description, as [start, end] seconds
    • Short, long and first-person (egocentric) video
    • Runs with Hugging Face transformers on one GPU
    • Apache-2.0, so commercial use is allowed

    Use Cases on Mixpeek

    Opening a search result at the exact seconds that match, after a retriever has found the right video
    Turning a highlight request into clip boundaries for an editor
    Labelling when an action happens in training or review footage

    Benchmarks

    DatasetMetricScoreSource
    Seven temporal grounding benchmarks (average)mIoU48.0Model card: MCG-NJU/TimeLens2-8B (self-reported, reported as state of the art on that suite)

    Performance

    Input SizeVideo sampled at 2 fps in the card's example, up to 480x480 per frame
    Embedding Dimn/a (outputs time spans as text)
    GPU LatencyInput dependent
    GPU ThroughputBatch dependent
    GPU MemoryModel dependent

    A generative model: it reads the whole clip for every query, so it is too slow to search a library on its own. Use it on the few clips a retriever returns. The card does not break the 48.0 average out by benchmark in text. We have not measured it.

    Frequently Asked Questions

    What does TimeLens2 do?

    It finds when something happens inside a video. Given a clip and a text description, it returns the time spans, in seconds, where the description is true.

    Can TimeLens2 search a whole video library?

    Not on its own. It reads the whole clip for each query, so it suits the handful of clips a search returns. Index the library with an embedding-based search first, then run TimeLens2 on the top results.

    Can I use TimeLens2 commercially?

    Yes. It is released under Apache-2.0.

    How do I use TimeLens2 with Mixpeek?

    Index the footage in a Mixpeek video collection and search it with a retriever, which returns scenes with timestamps. Run TimeLens2 on the top scene with the same query when you need tighter start and end times than the scene boundaries.

    Specification

    OrganizationMCG-NJU
    Retriever-
    Parameters8B
    LicenseApache-2.0
    Downloads/moN/A
    Likes17

    Research Paper

    TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs

    arxiv.org

    Build a pipeline with TimeLens2-8B

    Add this model to a processing pipeline alongside other extractors. Combine with retrieval stages for end-to-end search.

    Run it on your own data, free