NEWVectors or files. Pick a path.Start →
    Search & Discovery
    8 min read
    Updated 2026-10-11

    Why Does My Image Search Miss Small Objects in a Photo?

    Image search misses small objects because most systems store one vector per photo, computed from a copy shrunk to about 224 pixels. A 100-pixel object in a 4000x3000 photo becomes 5 to 8 pixels, smaller than one of the model's image patches, and CLIP's center crop drops a quarter of a wide photo before the model sees it. The fix is to embed crops of each photo as well as the whole, and rank each photo by its best-matching crop.

    Image Search
    Embeddings
    Object Detection
    CLIP
    SigLIP

    Why does my image search miss small objects in a photo?



    Most image search stores one vector per photo, and the model computes that vector from a copy of the photo shrunk to about 224 pixels on a side. A 100-pixel object in a 4000x3000 photo ends up 5 to 8 pixels wide, smaller than one of the 14- or 16-pixel patches the model reads the image in. The vector describes the whole scene, so a query for the small thing scores the photo low. The fix is to embed crops of each photo as well as the whole photo, and to rank each photo by its best-matching crop.

    What happens to a photo before the model sees it?



    Each model's preprocessing is in its published config on Hugging Face. The last column is arithmetic for a 4000x3000 photo with a 100-pixel-wide object in it.

    ModelWhat it does to the photoWhat the model readsA 100 px object in a 4000x3000 photo
    CLIP ViT-L/14Resizes the short side to 224 px, then center-crops 224x22416x16 patches of 14 pxAbout 7.5 px wide, and the crop drops 500 px from each side of the photo
    SigLIP base 224Resizes the whole photo to 224x224, squashing the aspect ratio14x14 patches of 16 pxAbout 5.6 px wide
    SigLIP 2 NaFlexKeeps the aspect ratio, up to 256 patches by defaultPatches of 16 px, each covering about 215x215 px of the photoAbout half of one patch
    Checked 2026-10-11 against each model's preprocessor_config.json.

    Why doesn't the vector capture what's in the corner?



    Two reasons stack up. The first is resolution: after the resize, a small object covers a fraction of one patch, so there is little signal for the model to encode. The second is training. CLIP-style models learn from image and caption pairs (Radford et al., 2021), and a caption describes the main subject of a picture. A photo of a warehouse aisle is captioned as a warehouse aisle, so the vector for that photo sits near "warehouse aisle" and far from "fire extinguisher", even when one hangs on the end of the shelf.

    The center crop adds a third failure for CLIP: on a 4:3 photo it removes a quarter of the width before the model sees anything, so an object near the left or right edge is gone entirely.

    Does a higher-resolution model fix it?



    It helps. SigLIP 2's NaFlex variant (Tschannen et al., 2025) keeps the aspect ratio, so nothing is cropped away, and it lets you raise the patch budget above the default 256. But the result is still one vector for the whole photo. A small object adds a little to that vector, and the rest of the scene adds much more, so a query about the small object still loses to photos where that object is the subject.

    How do I make image search find small objects?



    Give each region of the photo its own vector. There are two ways to choose the regions.

    ApproachHow it worksVectors per photoGood forWeak spot
    Whole photo onlyOne embedding of the resized photo1Queries about the scene or main subjectSmall and off-center objects
    TilingCut overlapping tiles at two or three grid sizes and embed each1 + 4 + 9 = 14 for a 2x2 and 3x3 gridAnything, including objects you can't name in advanceStorage and embedding cost grow about 14x
    Detect, then embedRun an open-vocabulary detector such as OWLv2 or Grounding DINO, crop each box, embed each cropOne per detected objectCatalogs of known object types: products, logos, vehicles, equipmentMisses what the detector was not asked to find
    Caption, then search textA vision language model writes a detailed description and you search the text1 text fieldQueries phrased in wordsDepends on the caption mentioning the object
    Keep the whole-photo vector in every case, so scene-level queries still work. At query time, score every crop and give each photo the score of its best crop. Return the crop's box with the result so the user sees where the match is.

    Tiles should overlap, or an object cut by a tile seam ends up half in two tiles and whole in none. The code below makes each tile one and a half times the size of its grid cell.
    import torch
    from PIL import Image
    from transformers import AutoModel, AutoProcessor
    
    model = AutoModel.from_pretrained("google/siglip-base-patch16-224")
    processor = AutoProcessor.from_pretrained("google/siglip-base-patch16-224")
    
    def tiles(img, grid):
        w, h = img.size
        tw, th = w / grid, h / grid
        for row in range(grid):
            for col in range(grid):
                cx, cy = (col + 0.5) * tw, (row + 0.5) * th
                # 1.5x the grid cell, so an object on a seam lands whole in some tile
                box = (max(0, cx - 0.75 * tw), max(0, cy - 0.75 * th),
                       min(w, cx + 0.75 * tw), min(h, cy + 0.75 * th))
                yield box, img.crop(box)
    
    def embed_photo(path):
        img = Image.open(path).convert("RGB")
        crops = [((0, 0) + img.size, img)]          # the whole photo
        for grid in (2, 3):
            crops += list(tiles(img, grid))         # 4 + 9 tiles
        inputs = processor(images=[c for _, c in crops], return_tensors="pt")
        with torch.no_grad():
            vecs = model.get_image_features(**inputs)
        return [box for box, _ in crops], torch.nn.functional.normalize(vecs, dim=-1)
    
    def score(query, boxes, vecs):
        inputs = processor(text=[query], padding="max_length", return_tensors="pt")
        with torch.no_grad():
            q = torch.nn.functional.normalize(model.get_text_features(**inputs), dim=-1)
        sims = (vecs @ q.T).squeeze(1)
        best = int(sims.argmax())
        return float(sims[best]), boxes[best]       # a photo scores as its best crop
    
    boxes, vecs = embed_photo("warehouse_aisle.jpg")
    print(score("a red fire extinguisher", boxes, vecs))
    In a vector database, store each crop as its own row with the photo's ID and the crop's box, search all rows, and keep the top row per photo.

    How much more does it cost?



    Tiling at 2x2 and 3x3 stores 14 vectors per photo where you stored 1, so storage and embedding compute grow about 14 times. Detection stores one vector per object found plus the whole photo, which is usually fewer, but you also run the detector on every photo. If most queries are about the scene, keep whole-photo search as the default and send queries that name a small object to the crop index.

    How do I do this with Mixpeek?



    Mixpeek reads your files where they already live, in S3, GCS or any bucket, and turns what's in them into data your software can search, classify and moderate. Its image extractor embeds each image with SigLIP base at 224x224, so on a whole photo it has the limit this guide describes. To search small objects, upload the crops or tiles as their own objects, record which photo each one came from, and add a deduplicate stage on that field to the retriever, so each photo comes back once, at the rank of its best crop. Prices are on the pricing page.

    For other approaches to the same problem, see mask-aware retrieval, which segments objects before searching, and open-vocabulary object detection. To find the exact same item rather than similar ones, see instance-level visual matching. To pick the embedding model, see the best multimodal embedding models and the best AI image search tools.

    Frequently Asked Questions



    Why does my image search miss small objects in a photo?



    It stores one vector per photo, computed from a copy shrunk to about 224 pixels. A small object becomes a few pixels, less than one of the model's patches, and the vector mostly describes the main subject. Embedding crops of the photo, and ranking each photo by its best crop, lets the small object match.

    Why does CLIP ignore objects at the edge of an image?



    CLIP's preprocessing resizes the short side to 224 pixels and then crops the center 224x224 square. On a 4:3 photo that removes a quarter of the width, so anything near the left or right edge never reaches the model.

    Should I tile my images or run an object detector first?



    Tile when you can't predict what people will search for, since tiling covers the whole photo. Detect first when the objects come from a known set, such as products or logos, because you store far fewer vectors. Either way, keep a whole-photo vector for scene-level queries.

    Does a bigger or higher-resolution embedding model solve it?



    Partly. SigLIP 2 NaFlex keeps the aspect ratio and can read more patches, so less detail is lost. It still makes one vector for the whole photo, and the rest of the scene outweighs a small object in it.
    Managed Mixpeek

    Put multimodal search to work

    Connect a bucket and Mixpeek runs the whole multimodal search pipeline for you: extraction, indexing, and search over your own objects. No models to wire up, nothing to host.

    Start with Managed
    MVS · bring your own

    Already have vectors?

    Keep your embeddings on your own cloud and run dense, sparse, and BM25 search directly on object storage. From $25/mo.

    Start with MVS

    Run this on your own data

    Point Mixpeek at the storage you already have and search your video, images, audio, and documents the way this guide describes. Build starts at $25/mo for up to 1M vectors.

    Search your own archiveRead Docs

    Related guides

    Search & Discovery

    How Do I Find Duplicate Photos, Including Edited and Resized Copies?

    A resized, recompressed, cropped or filtered photo has different bytes from the original, so a simple duplicate finder misses it. This explains the three ways to compare photos, what each one catches, the tools for a personal library and for an archive of millions, and how to set the threshold so you keep the right copy.

    Read guide →
    Search & Discovery

    How Do I Find Every Ad We've Run That Shows a Specific Product?

    Your ad archive has thousands of videos and you need every one where a particular product appears on screen, with the second it appears. File names and DAM tags rarely say that. This covers the methods that work, what each one misses, how to tell your product from a lookalike, and what it costs to index a whole archive.

    Read guide →
    Search & Discovery

    What Is Composite Clustering? Clustering Across Multiple Feature Spaces (and Clusters of Clusters)

    Two things people call composite clustering: multi-feature clustering groups documents once using several embedding spaces at once (text + image + audio + faces), with concatenate / independent / weighted combine strategies; composite (cluster-of-clusters) clustering groups the centroids of prior clusterings to reveal how your groupings relate. How each works, the normalization and per-feature-weight traps, how it differs from multi-vector retrieval, and the honest limits (composite is a pattern map, not a per-document cross-tab).

    Read guide →