NEWVectors or files. Pick a path.Start →
    Models/Detection & Recognition/google/owlv2-large-patch14-ensemble
    HFObject Detectionapache-2.0

    owlv2-large-patch14-ensemble

    by google

    Open-vocabulary OWLv2 detector for text-conditioned object search

    145Kdl/month
    46likes
    438Mparams
    Identifiers
    Model ID
    google/owlv2-large-patch14-ensemble
    Feature URI
    mixpeek://image_extractor@v1/google_owlv2_large_ensemble_v1

    Deploy owlv2-large-patch14-ensemble

    Single-tenant

    Mixpeek has no managed extractor for this model. On a single-tenant deployment you upload the weights and a custom plugin serves them next to the rest of your pipeline.

    Overview

    OWLv2 Large Patch14 Ensemble is Google's open-vocabulary detector for zero-shot object localization. It lets a pipeline search for objects described in text instead of relying only on a fixed supervised label set.

    On Mixpeek, OWLv2 is useful when an agent needs to find visual categories that change by task: a specific product shape, a UI control, damaged equipment, or a visual policy violation. The detector outputs boxes and labels that can be stored, filtered, and joined with embeddings or captions.

    Architecture

    Vision Transformer based open-vocabulary object detector. It aligns text queries and image regions so arbitrary text labels can guide detection at inference time.

    How it runs

    Inference INPUT MODEL OUTPUT Image object Text object owlv2-large-patch14-ensemble mp.inference Detections vector Vector store MVS owlv2-large-patch14-ensemble → embeddings, indexed for search
    owlv2-large-patch14-ensemble takes image and text, and Mixpeek indexes what it emits.

    Mixpeek SDK Integration

    // No extractor parameter takes a Hugging Face model id (checked against
    // GET /v1/discovery/extractors, which returns 13), so owlv2-large-patch14-ensemble runs
    // on your side and the output is upserted through POST
    // /v1/namespaces/{namespace_id}/documents/upsert. On Enterprise the other
    // path is to upload the weights instead: POST /v1/namespaces/{id}/models
    // accepts the huggingface format and a custom plugin loads them.
    const res = await fetch(
      "https://api.mixpeek.com/v1/namespaces/ns_your_namespace/documents/upsert",
      {
        method: "POST",
        headers: {
          Authorization: "Bearer API_KEY",
          "Content-Type": "application/json",
        },
        body: JSON.stringify({
          collection_id: "col_your_collection",
          documents: [
            {
              document_id: "asset-00412",
              // Boxes, masks, depth maps and anomaly scores are structured
              // results, not vectors. They go in payload and are reachable
              // through pre_filters on a retriever, not through similarity.
              payload: {
                detections: modelOutput,
                source_key: "archive/2026/asset-00412",
              },
            },
          ],
        }),
      },
    );
    
    // No managed alternative for an open label set. Two extractors do emit a
    // bbox, for the one thing each detects: document_graph_extractor@v1 per
    // layout block, face_identity_extractor@v1 per face. Nothing ships that
    // returns masks, depth maps or anomaly scores.

    Capabilities

    • Zero-shot object detection
    • Text-conditioned visual localization
    • Strong fit for dynamic agent queries
    • Apache 2.0 license

    Use Cases on Mixpeek

    Search frames for object classes not known during ingestion design
    Find UI controls or visual states from natural-language prompts
    Build long-tail visual filters over product and media libraries
    Pair open-vocabulary boxes with scene captions for agent evidence

    Specification

    FrameworkHF
    Organizationgoogle
    FeatureObject Detection
    Outputbbox + label
    Modalitiesvideo, image
    RetrieverObject Filter
    Parameters438M
    Licenseapache-2.0
    Downloads/mo145K
    Likes46

    Research Paper

    OWLv2 Large Patch14 Ensemble

    arxiv.org

    Build a pipeline with owlv2-large-patch14-ensemble

    Add this model to a processing pipeline alongside other extractors. Combine with retrieval stages for end-to-end search.

    Run it on your own data, free