NEWVectors or files. Pick a path.Start →
    Models/Alibaba-VELLDEPTH/CTVG-4B
    Unresolved: academic-use notice from TimeLens2, no commercial license granted

    CTVG-4B

    by Alibaba-VELLDEPTH

    A 4B model that returns the exact start and end time of a described moment in a video, with a confidence score for each interval

    Identifiers
    Model ID
    Alibaba-VELLDEPTH/CTVG-4B
    Feature URI

    Deploy CTVG-4B

    Single-tenant

    Mixpeek has no managed extractor for this model. On a single-tenant deployment you upload the weights and a custom plugin serves them next to the rest of your pipeline.

    Overview

    CTVG-4B is a video temporal grounding model from Alibaba, released in October 2026 with the paper Grounding with Confidence: Controllable Generative Video Temporal Grounding. Given a video and a text query such as "a person opens a door", it generates candidate time intervals and scores each one with a small confidence head, so you can rank the intervals, keep the best, or reject all of them when the event is not in the video.\n\nIt is a fine-tune of MCG-NJU/TimeLens2-4B, which uses the Qwen3-VL architecture. The card reports 67.39 [email protected] and 63.24 temporal IoU on the 320 queries of OMTG-Bench, and says most of that gain comes from non-maximum suppression rather than from the confidence head. Its license is unresolved: the card carries an academic-use notice from TimeLens2 and grants no commercial license.

    Architecture

    Qwen3-VL-architecture generator (about 4.8B parameters, merged BF16), fine-tuned from TimeLens2-4B with supervised training and then set-level reinforcement learning. A separate confidence head (LayerNorm, Linear 2560 to 256, GELU, Linear 256 to 1, sigmoid) scores each generated interval from the mean of its decoder states. Candidates pass temporal NMS at IoU 0.3 before the confidence threshold. Released inference settings: 2 fps, a 16,384-token context and at most 512 output tokens.

    Mixpeek SDK Integration

    # Index the videos first so search can find the right scenes. This uses
    # Mixpeek's multimodal extractor with scene splitting, as in the video
    # scene search recipe; CTVG-4B then runs on your side on the scenes it returns.
    import requests
    
    requests.post(
        "https://api.mixpeek.com/v1/buckets/bkt_your_bucket/objects",
        headers={"Authorization": "Bearer API_KEY", "X-Namespace": "ns_your_namespace"},
        json={"blobs": [{"property": "video", "type": "video",
                         "data": "s3://your-bucket/recordings/dock-cam-2.mp4"}]},
    )

    Capabilities

    • Returns start and end times for a text-described event in a video
    • Several candidate intervals per query, for events that repeat
    • A confidence score per interval, for ranking or rejecting
    • Thresholds can change without decoding the video again
    • Runs through the authors' Grounding-with-Confidence inference code

    Use Cases on Mixpeek

    Trimming a retrieved scene to the exact seconds a query describes
    Finding every time an action repeats in a recording
    Rejecting a search hit when the described event is not in the clip

    Benchmarks

    DatasetMetricScoreSource
    OMTG-Bench (320 queries)[email protected]67.39%Model card: Alibaba-VELLDEPTH/CTVG-4B (self-reported; threshold selected on the test set)
    OMTG-Bench (320 queries)Temporal IoU63.24%Model card (self-reported)
    OMTG-Bench (320 queries)EtF142.64%Model card (self-reported)

    Performance

    Input SizeA video file and a text query, read at 2 fps in the released configuration
    Embedding Dimn/a (outputs time intervals with confidence scores)
    GPU LatencyInput dependent
    GPU ThroughputBatch dependent
    GPU MemoryModel dependent

    The card says the operating threshold was chosen on the test set, so it is not a calibrated default, and that the confidence head's gain after NMS (+0.32 points) is not statistically established. It needs a CUDA GPU and the companion inference code. We have not measured it.

    Frequently Asked Questions

    What is CTVG-4B?

    A video temporal grounding model. You give it a video and a description of an event, and it returns the time intervals where that event happens, each with a confidence score.

    What is video temporal grounding?

    Finding when something happens in a video from a text description, as start and end times. Search tells you which video or scene matches; grounding tells you which seconds.

    Can I use CTVG-4B commercially?

    The card does not grant that. It preserves an academic-use notice from TimeLens2, its base model, and says public availability does not imply permission for commercial use. Check the upstream terms before using it in a product.

    How do I use it with Mixpeek?

    Index your videos with scene splitting so a retriever returns the matching scenes with their start and end times, then run CTVG-4B on those scenes with the same query to get the exact interval, as in the example on this page.

    Specification

    OrganizationAlibaba-VELLDEPTH
    Retriever-
    Parameters4.8B
    LicenseUnresolved: academic-use notice from TimeLens2, no commercial license granted
    Downloads/moN/A
    Likes3

    Research Paper

    Grounding with Confidence: Controllable Generative Video Temporal Grounding

    arxiv.org

    Build a pipeline with CTVG-4B

    Add this model to a processing pipeline alongside other extractors. Combine with retrieval stages for end-to-end search.

    Run it on your own data, free