siglip-so400m-patch14-384
by google
SigLIP SO400M — shape-optimized image-text encoder, a top-quality CLIP alternative
google/siglip-so400m-patch14-384Overview
SigLIP SO400M/14 at 384px pairs Google's sigmoid image-text loss with the SoViT-400M 'shape-optimized' backbone — a compute-efficient ViT that punches well above its 400M parameter count. It is one of the strongest open image-text encoders for zero-shot classification and retrieval, and a common default when teams want better accuracy than CLIP without going to billion-parameter models.
On Mixpeek, SigLIP SO400M is a visual embedding extractor for image and video-frame search, with fine-grained visual understanding that helps on detailed product, scene, and style queries.
Architecture
Shape-optimized ViT (SoViT-400M/14) image encoder at 384px with a paired text encoder, trained with the sigmoid (SigLIP) loss instead of softmax contrastive — which scales better and improves zero-shot accuracy at a given compute budget.
Key Capabilities
- •High-accuracy image+text embeddings (sigmoid loss)
- •Strong fine-grained visual retrieval and zero-shot classification
- •Efficient 400M shape-optimized backbone
- •384px input for detailed scenes and on-image text
Use Cases on Mixpeek
- •Fine-grained product and scene visual search
- •Zero-shot tagging where CLIP ViT-L is not accurate enough
- •Video keyframe embeddings for media retrieval
- •Cross-modal recall feeding a reranker
Tags
Use siglip-so400m-patch14-384 on Mixpeek
Build multimodal processing pipelines with this model and others. Extract features, run inference, and set up retrieval in Mixpeek Studio.
Open StudioHow It Runs on Mixpeek
On Mixpeek, siglip-so400m-patch14-384 runs as a managed extractor inside a processing pipeline. Point a bucket of zero shot image classification data at it, and Mixpeek handles GPU provisioning, batching, retries, and writing the outputs into a vector store you can query.
Extractor outputs land in the Mixpeek Vector Store (MVS), where you can combine them with retrieval, reranking, and filter stages to build end-to-end search and agent-perception pipelines, no model-serving infrastructure to maintain.
Explore on Mixpeek
Compare alternatives in this category
Hand-picked tools & platforms compared
Deep-dive technical guide
See how Mixpeek runs models as extractors
Store & search embeddings at scale
Usage-based pricing for pipelines
Compare models, APIs & infrastructure
Specification
View on HuggingFace
See model card, files, and community discussion