sam-vit-huge
by facebook
Promptable foundation model for image segmentation
facebook/sam-vit-hugemixpeek://image_extractor@v1/facebook_sam_vit_huge_v1Deploy sam-vit-huge
Single-tenantMixpeek has no managed extractor for this model. On a single-tenant deployment you upload the weights and a custom plugin serves them next to the rest of your pipeline.
Overview
SAM (Segment Anything Model) is Meta's foundation model for image segmentation. Given prompts like points, boxes, or text, it produces high-quality object masks. Trained on SA-1B: the largest segmentation dataset with 1 billion masks on 11M images.
On Mixpeek, SAM powers pixel-level object segmentation for precise content understanding, enabling mask-based filtering and region-specific feature extraction.
Architecture
ViT-H image encoder (632M params) with a lightweight mask decoder. Produces 256x256 low-res masks refined to full resolution. Supports multiple prompt types: points, boxes, and masks.
Mixpeek SDK Integration
// No extractor parameter takes a Hugging Face model id (checked against
// GET /v1/discovery/extractors, which returns 13), so sam-vit-huge runs
// on your side and the output is upserted through POST
// /v1/namespaces/{namespace_id}/documents/upsert. On Enterprise the other
// path is to upload the weights instead: POST /v1/namespaces/{id}/models
// accepts the huggingface format and a custom plugin loads them.
const res = await fetch(
"https://api.mixpeek.com/v1/namespaces/ns_your_namespace/documents/upsert",
{
method: "POST",
headers: {
Authorization: "Bearer API_KEY",
"Content-Type": "application/json",
},
body: JSON.stringify({
collection_id: "col_your_collection",
documents: [
{
document_id: "asset-00412",
// Boxes, masks, depth maps and anomaly scores are structured
// results, not vectors. They go in payload and are reachable
// through pre_filters on a retriever, not through similarity.
payload: {
detections: modelOutput,
source_key: "archive/2026/asset-00412",
},
},
],
}),
},
);
// No managed alternative for an open label set. Two extractors do emit a
// bbox, for the one thing each detects: document_graph_extractor@v1 per
// layout block, face_identity_extractor@v1 per face. Nothing ships that
// returns masks, depth maps or anomaly scores.Capabilities
- Promptable segmentation with points, boxes, or masks
- Automatic mask generation for everything in an image
- Zero-shot transfer competitive with supervised models
- Trained on 1 billion masks (SA-1B dataset)
Use Cases on Mixpeek
Benchmarks
| Dataset | Metric | Score | Source |
|---|---|---|---|
| SA-1B (segmentation) | mIoU | 79.3 | Kirillov et al., 2023: Table 1 |
| COCO (instance seg.) | AP | 46.5 | Kirillov et al., 2023: Table 7 |
Performance
Image encoder runs once; mask decoder runs per prompt (~6ms)
Common Pipeline Companions
Explore on Mixpeek
Compare alternatives in this category
Hand-picked tools & platforms compared
Deep-dive technical guide
See how Mixpeek runs models as extractors
Store & search embeddings at scale
Usage-based pricing for pipelines
Compare models, APIs & infrastructure
Specification
Research Paper
Segment Anything
arxiv.orgBuild a pipeline with sam-vit-huge
Add this model to a processing pipeline alongside other extractors. Combine with retrieval stages for end-to-end search.
Run it on your own data, free