Meta's SAM 3: Type a Phrase, Segment Every Instance
Summary
SAM 1 and 2 returned one mask per click. SAM 3 (Meta, November 2025, open weights) introduces promptable concept segmentation: a short noun phrase or exemplar image makes one model detect, segment, and track every instance of that concept across images and video. A shared Perception Encoder feeds a DETR-based detector and a SAM 2 memory tracker, trained on 4M+ auto-annotated concepts. Results: 55.7 cgF1 on SA-Co Gold vs 24.5 for OWLv2, 75-80% of human performance, 30ms per image with 100+ objects; SAM 3.1 tracks 16 objects in one pass at 32 fps. Already powers Instagram Edits and Marketplace View in Room. Explore: ai.meta.com/blog/segment-anything-model-3
About this video
SAM 1 and 2 returned one mask per click. SAM 3 (Meta, November 2025, open weights) introduces promptable concept segmentation: a short noun phrase or exemplar image makes one model detect, segment, and track every instance of that concept across images and video. A shared Perception Encoder feeds a DETR-based detector and a SAM 2 memory tracker, trained on 4M+ auto-annotated concepts. Results: 55.7 cgF1 on SA-Co Gold vs 24.5 for OWLv2, 75-80% of human performance, 30ms per image with 100+ objects; SAM 3.1 tracks 16 objects in one pass at 32 fps. Already powers Instagram Edits and Marketplace View in Room. Explore: ai.meta.com/blog/segment-anything-model-3