NEWVectors or files. Pick a path.Start →
    Back to Research

    Meta's SAM 3: Type a Phrase, Segment Every Instance

    SAM 3 introduces promptable concept segmentation: a short noun phrase or exemplar image makes one model detect, segment, and track every instance of that concept across images and video.

    Where this fits in Mixpeek

    Mixpeek decomposes video and images into searchable features — objects, regions, faces, on-screen text. Concept segmentation like SAM 3 is the extraction primitive behind that visual decomposition: instead of one mask per click, a text prompt surfaces every instance of a concept across a clip, which is exactly the granularity Mixpeek indexes so a query like 'yellow school bus' returns the timestamped moments it appears. It maps directly onto Mixpeek's object/region extractors feeding the multimodal warehouse.

    About this research

    SAM 1 and 2 returned one mask per click. SAM 3 (Meta, November 2025, open weights) introduces promptable concept segmentation: a short noun phrase or exemplar image makes one model detect, segment, and track every instance of that concept across images and video. A shared Perception Encoder feeds a DETR-based detector and a SAM 2 memory tracker, trained on 4M+ auto-annotated concepts. Results: 55.7 cgF1 on SA-Co Gold vs 24.5 for OWLv2, 75-80% of human performance, 30ms per image with 100+ objects; SAM 3.1 tracks 16 objects in one pass at 32 fps. Already powers Instagram Edits and Marketplace View in Room. Explore: ai.meta.com/blog/segment-anything-model-3

    sam3segmentationcomputer-visionvideo-understandingobject-trackingextraction

    Frequently asked questions

    Put the research to work

    Mixpeek turns video, images, audio, and documents in your object storage into searchable, timestamped results through one API — the retrieval stack these papers describe.

    Search your own data, free