NEWVectors or files. Pick a path.Start →
    Models/PekingU/rtdetr_v2_r18vd
    Apache-2.0

    rtdetr_v2_r18vd

    by PekingU

    Real-time object detection with no NMS step, at 20M parameters

    Identifiers
    Model ID
    PekingU/rtdetr_v2_r18vd
    Feature URI

    Deploy rtdetr_v2_r18vd

    Single-tenant

    Mixpeek has no managed extractor for this model. On a single-tenant deployment you upload the weights and a custom plugin serves them next to the rest of your pipeline.

    Overview

    RT-DETRv2 is a detection transformer built for real-time use, and the r18vd variant is the small one: 20.2 million parameters, 300 object queries, a 256-dimensional decoder. Because it is a DETR, it predicts a fixed set of objects directly and needs no non-maximum suppression pass, which removes a tuning knob and a source of nondeterminism from the pipeline.

    It is trained on COCO, so it detects COCO's 80 classes and nothing else. For a media pipeline that means people, vehicles, animals, furniture and common objects are covered, and your brand's product categories are not. An open-vocabulary detector is the right tool when the label set is yours rather than COCO's.

    Where it fits in a retrieval pipeline: detection output is metadata, not a vector. Boxes and labels become filterable fields on the document, so a query can ask for scenes containing a dog, and the ranking within those scenes comes from an embedding model.

    Architecture

    RT-DETRv2 detection transformer, ResNet-18 backbone with the vd stem, d_model 256, 1 encoder layer over the flattened feature map, 300 object queries, 20,209,716 parameters. Trained on COCO (80 classes). The r50vd sibling is the same architecture at 43,019,444 parameters. No NMS post-processing: the 300 queries are the predictions.

    Mixpeek SDK Integration

    // Detection output is metadata rather than a vector, so it belongs in the
    // document payload where pre_filters can reach it. Run the model on your side,
    // or upload the weights on a single-tenant deployment and have a plugin run it.
    const res = await fetch(
      "https://api.mixpeek.com/v1/namespaces/ns_your_namespace/documents/upsert",
      {
        method: "POST",
        headers: {
          Authorization: "Bearer API_KEY",
          "Content-Type": "application/json",
        },
        body: JSON.stringify({
          collection_id: "col_your_collection",
          documents: [
            {
              document_id: "scene-00412-0184",
              // The embedding comes from a visual encoder; this model supplies the
              // labels beside it.
              vectors: { "visual-embedding": yourVector },
              payload: {
                source_key: "archive/2026/ep-412.mp4",
                start_ms: 184000,
                detected: ["person", "dog", "bicycle"],
                detected_count: 3,
              },
            },
          ],
        }),
      },
    );

    Capabilities

    • 80 COCO classes with boxes and confidence scores
    • No non-maximum suppression, so no NMS threshold to tune
    • Small enough for frame-rate detection on modest GPUs
    • A 43M-parameter sibling (r50vd) when accuracy matters more than speed

    Use Cases on Mixpeek

    Tagging video frames with the objects present, as filterable metadata
    Counting people or vehicles per scene for archive analytics
    Gating an expensive extractor so it only runs on frames containing something
    Region proposals that a crop-and-embed step turns into searchable vectors

    Performance

    Input SizeSingle image or video frame
    GPU LatencyInput dependent
    GPU ThroughputBatch dependent
    GPU MemoryModel dependent

    Real-time throughput depends on the backbone, the input resolution and the GPU, and we have not measured it here. The paper publishes latency against COCO accuracy for each variant.

    Frequently Asked Questions

    What can RT-DETRv2 detect?

    The 80 COCO classes, because that is what this checkpoint is trained on. It cannot detect a category outside that list, and fine-tuning or an open-vocabulary model is the path when your labels are your own.

    Why is no NMS a big deal?

    Non-maximum suppression is a post-processing pass with its own IoU threshold, and that threshold changes your results without appearing in any model metric. A DETR predicts a fixed set of queries directly, so the detector's output is the detector's output.

    r18vd or r50vd?

    r18vd at 20.2M parameters is the one most people download, and r50vd at 43.0M trades throughput for accuracy. The architecture, the query count and the interface are identical, so swapping between them is a checkpoint change.

    Specification

    OrganizationPekingU
    Retriever-
    Parameters20M
    LicenseApache-2.0
    Downloads/moN/A
    Likes8

    Research Paper

    RT-DETRv2: Improved Baseline with Bag-of-Freebies for Real-Time Detection Transformer

    arxiv.org

    Build a pipeline with rtdetr_v2_r18vd

    Add this model to a processing pipeline alongside other extractors. Combine with retrieval stages for end-to-end search.

    Run it on your own data, free