NEWVectors or files. Pick a path.Start →
    Models/Segmentation/facebook/sam3.1
    HFSegmentationother

    sam3.1

    by facebook

    7x faster multi-object tracking via Object Multiplex shared-memory architecture

    77Kdl/month
    555likes
    848Mparams
    Identifiers
    Model ID
    facebook/sam3.1
    Feature URI
    mixpeek://image_extractor@v1/facebook_sam31_v1

    Overview

    SAM 3.1 is Meta's update to the Segment Anything Model 3 that introduces Object Multiplex, a shared-memory approach for joint multi-object tracking. Instead of processing each object independently, SAM 3.1 bundles all tracked objects into a single forward pass with global reasoning, delivering 7x faster inference at 128 objects on a single H100 GPU while improving accuracy on 6 of 7 video segmentation benchmarks.

    On Mixpeek, SAM 3.1 replaces SAM 3 as the default segmentation model for video analytics pipelines involving multiple simultaneous objects. The Object Multiplex architecture halves VRAM usage (8GB to 4GB FP16) while doubling throughput from 16 to 32 FPS, making multi-object tracking practical for production-scale video processing.

    Architecture

    Same detector-tracker architecture as SAM 3 (848M parameters) with Object Multiplex extension. Shared-memory joint processing of up to 16 objects per forward pass. DETR-based detector conditioned on text prompts, geometric prompts, and image exemplars. Global reasoning across all tracked objects simultaneously.

    Mixpeek SDK Integration

    // No extractor parameter takes a Hugging Face model id (checked against
    // GET /v1/discovery/extractors, which returns 13), so sam3.1 runs
    // on your side and the output is upserted through POST
    // /v1/namespaces/{namespace_id}/documents/upsert. On Enterprise the other
    // path is to upload the weights instead: POST /v1/namespaces/{id}/models
    // accepts the huggingface format and a custom plugin loads them.
    const res = await fetch(
      "https://api.mixpeek.com/v1/namespaces/ns_your_namespace/documents/upsert",
      {
        method: "POST",
        headers: {
          Authorization: "Bearer API_KEY",
          "Content-Type": "application/json",
        },
        body: JSON.stringify({
          collection_id: "col_your_collection",
          documents: [
            {
              document_id: "asset-00412",
              // Boxes, masks, depth maps and anomaly scores are structured
              // results, not vectors. They go in payload and are reachable
              // through pre_filters on a retriever, not through similarity.
              payload: {
                detections: modelOutput,
                source_key: "archive/2026/asset-00412",
              },
            },
          ],
        }),
      },
    );
    
    // No managed alternative for an open label set. Two extractors do emit a
    // bbox, for the one thing each detects: document_graph_extractor@v1 per
    // layout block, face_identity_extractor@v1 per face. Nothing ships that
    // returns masks, depth maps or anomaly scores.

    Capabilities

    • 7x faster than SAM 3 at 128 tracked objects (H100)
    • Object Multiplex: joint multi-object tracking in single forward pass
    • Improved on 6/7 VOS benchmarks including +2.0 on MOSEv2
    • Half the VRAM of SAM 3 (4GB vs 8GB FP16)
    • 32 FPS multi-object tracking on H100

    Use Cases on Mixpeek

    Multi-object video tracking at scale: people, products, vehicles across surveillance feeds
    Brand and logo tracking across advertising content with dozens of simultaneous objects
    Production video editing with real-time multi-object segmentation and masking

    Benchmarks

    DatasetMetricScoreSource
    MOSEv2 (video seg.)J&F+2.0 over SAM 3Meta, Mar 2026: SAM 3.1 Release
    SA-V (video seg.)J&FImproved on 6/7 benchmarksMeta, Mar 2026: SAM 3.1 Release

    Performance

    Input Size1024x1024 px
    GPU Latency~31ms / frame (H100, 16 objects multiplex)
    GPU Throughput~32 FPS (H100, multi-object)
    GPU Memory~4 GB (FP16, half of SAM 3)

    Specification

    FrameworkHF
    Organizationfacebook
    FeatureSegmentation
    Outputmask + label
    Modalitiesvideo, image
    RetrieverMask Filter
    Parameters848M
    Licenseother
    Downloads/mo77K
    Likes555

    Research Paper

    SAM 3.1: Faster Multi-Object Tracking with Object Multiplex

    arxiv.org

    Build a pipeline with sam3.1

    Add this model to a processing pipeline alongside other extractors. Combine with retrieval stages for end-to-end search.

    Run it on your own data, free