NEWVectors or files. Pick a path.Start →
    Search & Discovery
    8 min read
    Updated 2026-09-29

    How Do I Find Duplicate Photos, Including Edited and Resized Copies?

    A resized, recompressed, cropped or filtered photo has different bytes from the original, so a simple duplicate finder misses it. This explains the three ways to compare photos, what each one catches, the tools for a personal library and for an archive of millions, and how to set the threshold so you keep the right copy.

    Duplicate Photos
    Image Deduplication
    Perceptual Hashing
    Image Search
    Digital Asset Management

    How do I find duplicate photos, including edited copies?



    Compare the photos by what they look like as well as by their bytes. A file checksum finds exact copies. A perceptual hash, a short fingerprint of the picture's overall shape and brightness, finds the same photo after resizing, recompression or a format change. An image embedding from a vision model finds copies that were cropped, recoloured, filtered or had text added. Run them in that order, cheapest first, group the matches, and keep the best copy from each group.

    Why doesn't my duplicate finder catch edited copies?



    Most duplicate finders compare file contents or checksums, so they only match files that are identical byte for byte. Saving a photo at a different size or quality, exporting it from an editor, or sending it through a messaging app rewrites every byte while the picture looks the same. To a checksum those are unrelated files.

    What are the ways to compare photos?



    MethodCatchesMissesCost
    File checksum (for example SHA-256)Exact byte-for-byte copies, whatever the file nameAny re-save, resize or format changeVery low
    Perceptual hash (aHash, dHash, pHash)Resized, recompressed and format-converted copies; small colour changesCrops, rotations, mirrored copies, heavy filters or overlaysLow
    Image embedding from a vision modelCrops, colour grades, filters, overlaid text, and the same scene shot a second laterCan group photos that are similar and distinct, such as a burst of near-identical shotsModerate, usually a GPU for large sets
    The embedding method is the most forgiving and the one that needs a threshold chosen with care, because "the same photo edited" and "a different photo of the same thing" can sit close together. Self-supervised models such as DINOv2 tend to separate those two cases better than models trained to match images with text.

    What tools can I use?



  1. For a personal library: Apple Photos on recent iOS and macOS versions has a Duplicates album
  2. that finds exact and near-identical photos and merges them. On a computer, open-source tools such as dupeGuru (picture mode) and Czkawka (similar images) compare folders with perceptual matching.
  3. For a shared drive or a DAM: many digital asset management systems flag exact duplicates on
  4. upload; fewer catch edited copies, so check before assuming yours does.
  5. For an archive of millions: compute checksums and perceptual hashes for everything, then
  6. embeddings, and use a vector index so each photo is compared only with its nearest neighbours instead of with every other photo. Comparing every pair grows with the square of the collection and stops being practical long before a million images.

    How do I choose the threshold and which copy to keep?



    Take a few hundred photos you know well, label which pairs are true duplicates, and look at the similarity scores. Set the cut-off where true duplicates and distinct photos separate, and send the borderline band to a person rather than deleting it automatically. Then keep one copy per group by a rule you can state: the largest resolution, the earliest date, or the one with the most metadata, since edited copies often lose their camera data and location.

    Never delete in the same step that finds duplicates. Move them aside, check a sample, then delete.

    How do I do this with Mixpeek?



    Point Mixpeek at the bucket that holds the photos. Every object is fingerprinted with a SHA-256 content hash on the way in, so exact copies are recognised at ingest and not processed twice. The image extractor embeds each photo, and a retriever finds each photo's near matches by similarity, while clustering groups them across the whole archive and the deduplicate stage removes repeats from search results. The photos stay in your own storage. At the published rate of $1.50 per thousand images, embedding 100,000 photos costs about $150 in processing.

    Related: why image search returns lookalikes instead of the exact item, the best image similarity search tools, the best reverse image search APIs and the best video deduplication tools.

    Frequently Asked Questions



    What is the difference between a duplicate and a near-duplicate photo?



    A duplicate is the same file, byte for byte. A near-duplicate is the same picture saved differently: resized, recompressed, cropped, filtered or converted to another format. Checksums find duplicates; perceptual hashes and image embeddings find near-duplicates.

    What is a perceptual hash?



    A short fingerprint computed from a shrunken, greyscale version of the picture, so it stays almost the same when the photo is resized or recompressed. Two photos are compared by counting how many bits of their fingerprints differ; a small difference means the same picture.

    Can I find duplicates that were cropped or had a filter applied?



    Yes, with image embeddings from a vision model, which compare the content of the pictures. Perceptual hashes usually miss crops and heavy filters. Check a sample of matches, because embeddings can also group photos that are similar and not duplicates.

    How do I find duplicates in a very large photo archive?



    Avoid comparing every photo with every other one. Group exact copies by checksum first, then index perceptual hashes or embeddings in a nearest-neighbour index so each photo is compared only with its closest matches, and review the borderline pairs by hand before removing anything.
    Managed Mixpeek

    Put multimodal search to work

    Connect a bucket and Mixpeek runs the whole multimodal search pipeline for you: extraction, indexing, and search over your own objects. No models to wire up, nothing to host.

    Start with Managed
    MVS · bring your own

    Already have vectors?

    Keep your embeddings on your own cloud and run dense, sparse, and BM25 search directly on object storage. From $25/mo.

    Start with MVS

    Run this on your own data

    Point Mixpeek at the storage you already have and search your video, images, audio, and documents the way this guide describes. Build starts at $25/mo for up to 1M vectors.

    Search your own archiveRead Docs