NEWVectors or files. Pick a path.Start →

    Image Deduplication Pipeline

    Identify near-duplicate and visually similar images in your collection using embedding-based clustering. Groups images by visual similarity, flags duplicates, and provides confidence scores to help clean up redundant content in media libraries and product catalogs.

    image
    Multi-Stage
    import requests
    from mixpeek import Mixpeek
    client = Mixpeek(api_key="YOUR_API_KEY", namespace="image-library")
    API = "https://api.mixpeek.com/v1"
    HEADERS = {"Authorization": "Bearer YOUR_API_KEY", "X-Namespace": "image-library"}
    # 1. A bucket for the library, and a collection that embeds each image
    bucket = client.buckets.create(
    bucket_name="image-library",
    bucket_schema={"properties": {"image": {"type": "image"}}},
    )
    collection = client.collections.create(
    collection_name="image_library",
    source={"type": "bucket", "bucket_ids": [bucket["bucket_id"]]},
    feature_extractor={
    "feature_extractor_name": "multimodal_extractor",
    "version": "v1",
    },
    )
    client.buckets.upload(
    bucket["bucket_id"],
    blobs=[{"property": "image", "type": "image", "data": "s3://your-bucket/images/IMG_0001.jpg"}],
    )
    client.collections.trigger(collection["collection_id"])
    # 2. HDBSCAN with a minimum group size of 2 turns near-duplicates into small,
    # tight clusters. The SDK has no clusters resource, so this is REST.
    cluster = requests.post(API + "/clusters", headers=HEADERS, json={
    "cluster_name": "near-duplicates",
    "collection_ids": [collection["collection_id"]],
    "cluster_type": "vector",
    "vector_config": {
    "feature_uris": ["mixpeek://multimodal_extractor@v1/vertex_multimodal_embedding"],
    "clustering_method": "hdbscan",
    "algorithm_params": {"min_cluster_size": 2},
    },
    }).json()
    requests.post(API + "/clusters/" + cluster["cluster_id"] + "/execute", headers=HEADERS, json={})
    # 3. Once the execution finishes, review the groups
    groups = requests.get(API + "/clusters/" + cluster["cluster_id"] + "/groups", headers=HEADERS).json()
    print(groups["total_groups"], "groups of near-duplicates")

    Feature Extractors

    Multimodal Extractor

    Unified embeddings for video, audio, image, and text: scene/silence chunking, Whisper transcription, thumbnails, and Gemini vision.

    Retriever Stages