NEWVectors or files. Pick a path.Start →
    Models/Zero Shot Image Classification/google/siglip-so400m-patch14-384
    Zero Shot Image Classificationtransformersapache-2.0

    siglip-so400m-patch14-384

    by google

    SigLIP SO400M — shape-optimized image-text encoder, a top-quality CLIP alternative

    Identifier
    Model ID
    google/siglip-so400m-patch14-384

    Overview

    SigLIP SO400M/14 at 384px pairs Google's sigmoid image-text loss with the SoViT-400M 'shape-optimized' backbone — a compute-efficient ViT that punches well above its 400M parameter count. It is one of the strongest open image-text encoders for zero-shot classification and retrieval, and a common default when teams want better accuracy than CLIP without going to billion-parameter models.

    On Mixpeek, SigLIP SO400M is a visual embedding extractor for image and video-frame search, with fine-grained visual understanding that helps on detailed product, scene, and style queries.

    Architecture

    Shape-optimized ViT (SoViT-400M/14) image encoder at 384px with a paired text encoder, trained with the sigmoid (SigLIP) loss instead of softmax contrastive — which scales better and improves zero-shot accuracy at a given compute budget.

    Key Capabilities

    • High-accuracy image+text embeddings (sigmoid loss)
    • Strong fine-grained visual retrieval and zero-shot classification
    • Efficient 400M shape-optimized backbone
    • 384px input for detailed scenes and on-image text

    Use Cases on Mixpeek

    • Fine-grained product and scene visual search
    • Zero-shot tagging where CLIP ViT-L is not accurate enough
    • Video keyframe embeddings for media retrieval
    • Cross-modal recall feeding a reranker

    Tags

    transformerssafetensorssiglipzero-shot-image-classificationvisionarxiv:2303.15343arxiv:2305.13035arxiv:2209.06794license:apache-2.0endpoints_compatibleregion:us

    Use siglip-so400m-patch14-384 on Mixpeek

    Build multimodal processing pipelines with this model and others. Extract features, run inference, and set up retrieval in Mixpeek Studio.

    Open Studio

    How It Runs on Mixpeek

    On Mixpeek, siglip-so400m-patch14-384 runs as a managed extractor inside a processing pipeline. Point a bucket of zero shot image classification data at it, and Mixpeek handles GPU provisioning, batching, retries, and writing the outputs into a vector store you can query.

    Extractor outputs land in the Mixpeek Vector Store (MVS), where you can combine them with retrieval, reranking, and filter stages to build end-to-end search and agent-perception pipelines, no model-serving infrastructure to maintain.