How do I find duplicate photos, including edited copies?
Compare the photos by what they look like as well as by their bytes. A file checksum finds exact copies. A perceptual hash, a short fingerprint of the picture's overall shape and brightness, finds the same photo after resizing, recompression or a format change. An image embedding from a vision model finds copies that were cropped, recoloured, filtered or had text added. Run them in that order, cheapest first, group the matches, and keep the best copy from each group.
Why doesn't my duplicate finder catch edited copies?
Most duplicate finders compare file contents or checksums, so they only match files that are identical byte for byte. Saving a photo at a different size or quality, exporting it from an editor, or sending it through a messaging app rewrites every byte while the picture looks the same. To a checksum those are unrelated files.
What are the ways to compare photos?
| Method | Catches | Misses | Cost |
| File checksum (for example SHA-256) | Exact byte-for-byte copies, whatever the file name | Any re-save, resize or format change | Very low |
| Perceptual hash (aHash, dHash, pHash) | Resized, recompressed and format-converted copies; small colour changes | Crops, rotations, mirrored copies, heavy filters or overlays | Low |
| Image embedding from a vision model | Crops, colour grades, filters, overlaid text, and the same scene shot a second later | Can group photos that are similar and distinct, such as a burst of near-identical shots | Moderate, usually a GPU for large sets |
What tools can I use?
How do I choose the threshold and which copy to keep?
Take a few hundred photos you know well, label which pairs are true duplicates, and look at the similarity scores. Set the cut-off where true duplicates and distinct photos separate, and send the borderline band to a person rather than deleting it automatically. Then keep one copy per group by a rule you can state: the largest resolution, the earliest date, or the one with the most metadata, since edited copies often lose their camera data and location.
Never delete in the same step that finds duplicates. Move them aside, check a sample, then delete.
How do I do this with Mixpeek?
Point Mixpeek at the bucket that holds the photos. Every object is fingerprinted with a SHA-256 content hash on the way in, so exact copies are recognised at ingest and not processed twice. The image extractor embeds each photo, and a retriever finds each photo's near matches by similarity, while clustering groups them across the whole archive and the deduplicate stage removes repeats from search results. The photos stay in your own storage. At the published rate of $1.50 per thousand images, embedding 100,000 photos costs about $150 in processing.
Related: why image search returns lookalikes instead of the exact item, the best image similarity search tools, the best reverse image search APIs and the best video deduplication tools.
Frequently Asked Questions
What is the difference between a duplicate and a near-duplicate photo?
A duplicate is the same file, byte for byte. A near-duplicate is the same picture saved differently: resized, recompressed, cropped, filtered or converted to another format. Checksums find duplicates; perceptual hashes and image embeddings find near-duplicates.
What is a perceptual hash?
A short fingerprint computed from a shrunken, greyscale version of the picture, so it stays almost the same when the photo is resized or recompressed. Two photos are compared by counting how many bits of their fingerprints differ; a small difference means the same picture.
Can I find duplicates that were cropped or had a filter applied?
Yes, with image embeddings from a vision model, which compare the content of the pictures. Perceptual hashes usually miss crops and heavy filters. Check a sample of matches, because embeddings can also group photos that are similar and not duplicates.
How do I find duplicates in a very large photo archive?
Avoid comparing every photo with every other one. Group exact copies by checksum first, then index perceptual hashes or embeddings in a nearest-neighbour index so each photo is compared only with its closest matches, and review the borderline pairs by hand before removing anything.