yolos-tiny
by hustvl
You Only Look at One Sequence, ViT-based real-time object detection
hustvl/yolos-tinymixpeek://image_extractor@v1/hustvl_yolos_tiny_v1Deploy yolos-tiny
Single-tenantMixpeek has no managed extractor for this model. On a single-tenant deployment you upload the weights and a custom plugin serves them next to the rest of your pipeline.
Overview
YOLOS adapts the Vision Transformer (ViT) architecture for object detection by simply appending detection tokens to the input sequence. It demonstrates that a pure transformer can perform object detection without any convolutional components.
On Mixpeek, YOLOS Tiny provides a lightweight, fast alternative to DETR for object detection tasks where speed is prioritized over maximum accuracy.
Architecture
Vision Transformer (ViT-Tiny) with 12 layers. Appends 100 learnable detection tokens to the image patch sequence. Uses bipartite matching loss like DETR.
Mixpeek SDK Integration
// No extractor parameter takes a Hugging Face model id (checked against
// GET /v1/discovery/extractors, which returns 13), so yolos-tiny runs
// on your side and the output is upserted through POST
// /v1/namespaces/{namespace_id}/documents/upsert. On Enterprise the other
// path is to upload the weights instead: POST /v1/namespaces/{id}/models
// accepts the huggingface format and a custom plugin loads them.
const res = await fetch(
"https://api.mixpeek.com/v1/namespaces/ns_your_namespace/documents/upsert",
{
method: "POST",
headers: {
Authorization: "Bearer API_KEY",
"Content-Type": "application/json",
},
body: JSON.stringify({
collection_id: "col_your_collection",
documents: [
{
document_id: "asset-00412",
// Boxes, masks, depth maps and anomaly scores are structured
// results, not vectors. They go in payload and are reachable
// through pre_filters on a retriever, not through similarity.
payload: {
detections: modelOutput,
source_key: "archive/2026/asset-00412",
},
},
],
}),
},
);
// No managed alternative for an open label set. Two extractors do emit a
// bbox, for the one thing each detects: document_graph_extractor@v1 per
// layout block, face_identity_extractor@v1 per face. Nothing ships that
// returns masks, depth maps or anomaly scores.Capabilities
- Lightweight ViT-based object detection
- Fast inference suitable for real-time processing
- COCO object categories
- Pure transformer architecture (no CNN backbone)
Use Cases on Mixpeek
Benchmarks
| Dataset | Metric | Score | Source |
|---|---|---|---|
| COCO val2017 | AP (box) | 30.4 | Fang et al., 2021: Table 1 |
| COCO val2017 | AP50 | 48.6 | Fang et al., 2021: Table 1 |
Performance
6.5M params: optimized for edge and high-throughput scenarios
Common Pipeline Companions
Explore on Mixpeek
Compare alternatives in this category
Hand-picked tools & platforms compared
Deep-dive technical guide
See how Mixpeek runs models as extractors
Store & search embeddings at scale
Usage-based pricing for pipelines
Compare models, APIs & infrastructure
Specification
Research Paper
You Only Look at One Sequence
arxiv.orgBuild a pipeline with yolos-tiny
Add this model to a processing pipeline alongside other extractors. Combine with retrieval stages for end-to-end search.
Run it on your own data, free