NEWVectors or files. Pick a path.Start →
    HFObject Detectionapache-2.0

    yolos-tiny

    by hustvl

    You Only Look at One Sequence, ViT-based real-time object detection

    158Kdl/month
    282likes
    6Mparams
    Identifiers
    Model ID
    hustvl/yolos-tiny
    Feature URI
    mixpeek://image_extractor@v1/hustvl_yolos_tiny_v1

    Deploy yolos-tiny

    Single-tenant

    Mixpeek has no managed extractor for this model. On a single-tenant deployment you upload the weights and a custom plugin serves them next to the rest of your pipeline.

    Overview

    YOLOS adapts the Vision Transformer (ViT) architecture for object detection by simply appending detection tokens to the input sequence. It demonstrates that a pure transformer can perform object detection without any convolutional components.

    On Mixpeek, YOLOS Tiny provides a lightweight, fast alternative to DETR for object detection tasks where speed is prioritized over maximum accuracy.

    Architecture

    Vision Transformer (ViT-Tiny) with 12 layers. Appends 100 learnable detection tokens to the image patch sequence. Uses bipartite matching loss like DETR.

    Mixpeek SDK Integration

    // No extractor parameter takes a Hugging Face model id (checked against
    // GET /v1/discovery/extractors, which returns 13), so yolos-tiny runs
    // on your side and the output is upserted through POST
    // /v1/namespaces/{namespace_id}/documents/upsert. On Enterprise the other
    // path is to upload the weights instead: POST /v1/namespaces/{id}/models
    // accepts the huggingface format and a custom plugin loads them.
    const res = await fetch(
      "https://api.mixpeek.com/v1/namespaces/ns_your_namespace/documents/upsert",
      {
        method: "POST",
        headers: {
          Authorization: "Bearer API_KEY",
          "Content-Type": "application/json",
        },
        body: JSON.stringify({
          collection_id: "col_your_collection",
          documents: [
            {
              document_id: "asset-00412",
              // Boxes, masks, depth maps and anomaly scores are structured
              // results, not vectors. They go in payload and are reachable
              // through pre_filters on a retriever, not through similarity.
              payload: {
                detections: modelOutput,
                source_key: "archive/2026/asset-00412",
              },
            },
          ],
        }),
      },
    );
    
    // No managed alternative for an open label set. Two extractors do emit a
    // bbox, for the one thing each detects: document_graph_extractor@v1 per
    // layout block, face_identity_extractor@v1 per face. Nothing ships that
    // returns masks, depth maps or anomaly scores.

    Capabilities

    • Lightweight ViT-based object detection
    • Fast inference suitable for real-time processing
    • COCO object categories
    • Pure transformer architecture (no CNN backbone)

    Use Cases on Mixpeek

    Real-time video analysis where low latency is critical
    Edge deployment scenarios with limited compute
    High-throughput batch processing of large video archives

    Benchmarks

    DatasetMetricScoreSource
    COCO val2017AP (box)30.4Fang et al., 2021: Table 1
    COCO val2017AP5048.6Fang et al., 2021: Table 1

    Performance

    Input Size512×864 px
    GPU Latency~6ms / image (A100)
    CPU Latency~55ms / image
    GPU Throughput~165 images/sec (A100)
    GPU Memory~0.4 GB

    6.5M params: optimized for edge and high-throughput scenarios

    Specification

    FrameworkHF
    Organizationhustvl
    FeatureObject Detection
    Outputbbox + label
    Modalitiesvideo, image
    RetrieverObject Filter
    Parameters6M
    Licenseapache-2.0
    Downloads/mo158K
    Likes282

    Research Paper

    You Only Look at One Sequence

    arxiv.org

    Build a pipeline with yolos-tiny

    Add this model to a processing pipeline alongside other extractors. Combine with retrieval stages for end-to-end search.

    Run it on your own data, free