rtdetr_v2_r18vd
by PekingU
Real-time object detection with no NMS step, at 20M parameters
PekingU/rtdetr_v2_r18vdDeploy rtdetr_v2_r18vd
Single-tenantMixpeek has no managed extractor for this model. On a single-tenant deployment you upload the weights and a custom plugin serves them next to the rest of your pipeline.
Overview
RT-DETRv2 is a detection transformer built for real-time use, and the r18vd variant is the small one: 20.2 million parameters, 300 object queries, a 256-dimensional decoder. Because it is a DETR, it predicts a fixed set of objects directly and needs no non-maximum suppression pass, which removes a tuning knob and a source of nondeterminism from the pipeline.
It is trained on COCO, so it detects COCO's 80 classes and nothing else. For a media pipeline that means people, vehicles, animals, furniture and common objects are covered, and your brand's product categories are not. An open-vocabulary detector is the right tool when the label set is yours rather than COCO's.
Where it fits in a retrieval pipeline: detection output is metadata, not a vector. Boxes and labels become filterable fields on the document, so a query can ask for scenes containing a dog, and the ranking within those scenes comes from an embedding model.
Architecture
RT-DETRv2 detection transformer, ResNet-18 backbone with the vd stem, d_model 256, 1 encoder layer over the flattened feature map, 300 object queries, 20,209,716 parameters. Trained on COCO (80 classes). The r50vd sibling is the same architecture at 43,019,444 parameters. No NMS post-processing: the 300 queries are the predictions.
Mixpeek SDK Integration
// Detection output is metadata rather than a vector, so it belongs in the
// document payload where pre_filters can reach it. Run the model on your side,
// or upload the weights on a single-tenant deployment and have a plugin run it.
const res = await fetch(
"https://api.mixpeek.com/v1/namespaces/ns_your_namespace/documents/upsert",
{
method: "POST",
headers: {
Authorization: "Bearer API_KEY",
"Content-Type": "application/json",
},
body: JSON.stringify({
collection_id: "col_your_collection",
documents: [
{
document_id: "scene-00412-0184",
// The embedding comes from a visual encoder; this model supplies the
// labels beside it.
vectors: { "visual-embedding": yourVector },
payload: {
source_key: "archive/2026/ep-412.mp4",
start_ms: 184000,
detected: ["person", "dog", "bicycle"],
detected_count: 3,
},
},
],
}),
},
);Capabilities
- 80 COCO classes with boxes and confidence scores
- No non-maximum suppression, so no NMS threshold to tune
- Small enough for frame-rate detection on modest GPUs
- A 43M-parameter sibling (r50vd) when accuracy matters more than speed
Use Cases on Mixpeek
Performance
Real-time throughput depends on the backbone, the input resolution and the GPU, and we have not measured it here. The paper publishes latency against COCO accuracy for each variant.
Common Pipeline Companions
Frequently Asked Questions
What can RT-DETRv2 detect?
The 80 COCO classes, because that is what this checkpoint is trained on. It cannot detect a category outside that list, and fine-tuning or an open-vocabulary model is the path when your labels are your own.
Why is no NMS a big deal?
Non-maximum suppression is a post-processing pass with its own IoU threshold, and that threshold changes your results without appearing in any model metric. A DETR predicts a fixed set of queries directly, so the detector's output is the detector's output.
r18vd or r50vd?
r18vd at 20.2M parameters is the one most people download, and r50vd at 43.0M trades throughput for accuracy. The architecture, the query count and the interface are identical, so swapping between them is a checkpoint change.
Specification
Research Paper
RT-DETRv2: Improved Baseline with Bag-of-Freebies for Real-Time Detection Transformer
arxiv.orgBuild a pipeline with rtdetr_v2_r18vd
Add this model to a processing pipeline alongside other extractors. Combine with retrieval stages for end-to-end search.
Run it on your own data, free