Why does my image search miss small objects in a photo?
Most image search stores one vector per photo, and the model computes that vector from a copy of the photo shrunk to about 224 pixels on a side. A 100-pixel object in a 4000x3000 photo ends up 5 to 8 pixels wide, smaller than one of the 14- or 16-pixel patches the model reads the image in. The vector describes the whole scene, so a query for the small thing scores the photo low. The fix is to embed crops of each photo as well as the whole photo, and to rank each photo by its best-matching crop.
What happens to a photo before the model sees it?
Each model's preprocessing is in its published config on Hugging Face. The last column is arithmetic for a 4000x3000 photo with a 100-pixel-wide object in it.
| Model | What it does to the photo | What the model reads | A 100 px object in a 4000x3000 photo |
| CLIP ViT-L/14 | Resizes the short side to 224 px, then center-crops 224x224 | 16x16 patches of 14 px | About 7.5 px wide, and the crop drops 500 px from each side of the photo |
| SigLIP base 224 | Resizes the whole photo to 224x224, squashing the aspect ratio | 14x14 patches of 16 px | About 5.6 px wide |
| SigLIP 2 NaFlex | Keeps the aspect ratio, up to 256 patches by default | Patches of 16 px, each covering about 215x215 px of the photo | About half of one patch |
Why doesn't the vector capture what's in the corner?
Two reasons stack up. The first is resolution: after the resize, a small object covers a fraction of one patch, so there is little signal for the model to encode. The second is training. CLIP-style models learn from image and caption pairs (Radford et al., 2021), and a caption describes the main subject of a picture. A photo of a warehouse aisle is captioned as a warehouse aisle, so the vector for that photo sits near "warehouse aisle" and far from "fire extinguisher", even when one hangs on the end of the shelf.
The center crop adds a third failure for CLIP: on a 4:3 photo it removes a quarter of the width before the model sees anything, so an object near the left or right edge is gone entirely.
Does a higher-resolution model fix it?
It helps. SigLIP 2's NaFlex variant (Tschannen et al., 2025) keeps the aspect ratio, so nothing is cropped away, and it lets you raise the patch budget above the default 256. But the result is still one vector for the whole photo. A small object adds a little to that vector, and the rest of the scene adds much more, so a query about the small object still loses to photos where that object is the subject.
How do I make image search find small objects?
Give each region of the photo its own vector. There are two ways to choose the regions.
| Approach | How it works | Vectors per photo | Good for | Weak spot |
| Whole photo only | One embedding of the resized photo | 1 | Queries about the scene or main subject | Small and off-center objects |
| Tiling | Cut overlapping tiles at two or three grid sizes and embed each | 1 + 4 + 9 = 14 for a 2x2 and 3x3 grid | Anything, including objects you can't name in advance | Storage and embedding cost grow about 14x |
| Detect, then embed | Run an open-vocabulary detector such as OWLv2 or Grounding DINO, crop each box, embed each crop | One per detected object | Catalogs of known object types: products, logos, vehicles, equipment | Misses what the detector was not asked to find |
| Caption, then search text | A vision language model writes a detailed description and you search the text | 1 text field | Queries phrased in words | Depends on the caption mentioning the object |
Tiles should overlap, or an object cut by a tile seam ends up half in two tiles and whole in none. The code below makes each tile one and a half times the size of its grid cell.
import torch
from PIL import Image
from transformers import AutoModel, AutoProcessor
model = AutoModel.from_pretrained("google/siglip-base-patch16-224")
processor = AutoProcessor.from_pretrained("google/siglip-base-patch16-224")
def tiles(img, grid):
w, h = img.size
tw, th = w / grid, h / grid
for row in range(grid):
for col in range(grid):
cx, cy = (col + 0.5) * tw, (row + 0.5) * th
# 1.5x the grid cell, so an object on a seam lands whole in some tile
box = (max(0, cx - 0.75 * tw), max(0, cy - 0.75 * th),
min(w, cx + 0.75 * tw), min(h, cy + 0.75 * th))
yield box, img.crop(box)
def embed_photo(path):
img = Image.open(path).convert("RGB")
crops = [((0, 0) + img.size, img)] # the whole photo
for grid in (2, 3):
crops += list(tiles(img, grid)) # 4 + 9 tiles
inputs = processor(images=[c for _, c in crops], return_tensors="pt")
with torch.no_grad():
vecs = model.get_image_features(**inputs)
return [box for box, _ in crops], torch.nn.functional.normalize(vecs, dim=-1)
def score(query, boxes, vecs):
inputs = processor(text=[query], padding="max_length", return_tensors="pt")
with torch.no_grad():
q = torch.nn.functional.normalize(model.get_text_features(**inputs), dim=-1)
sims = (vecs @ q.T).squeeze(1)
best = int(sims.argmax())
return float(sims[best]), boxes[best] # a photo scores as its best crop
boxes, vecs = embed_photo("warehouse_aisle.jpg")
print(score("a red fire extinguisher", boxes, vecs))How much more does it cost?
Tiling at 2x2 and 3x3 stores 14 vectors per photo where you stored 1, so storage and embedding compute grow about 14 times. Detection stores one vector per object found plus the whole photo, which is usually fewer, but you also run the detector on every photo. If most queries are about the scene, keep whole-photo search as the default and send queries that name a small object to the crop index.
How do I do this with Mixpeek?
Mixpeek reads your files where they already live, in S3, GCS or any bucket, and turns what's in them into data your software can search, classify and moderate. Its image extractor embeds each image with SigLIP base at 224x224, so on a whole photo it has the limit this guide describes. To search small objects, upload the crops or tiles as their own objects, record which photo each one came from, and add a deduplicate stage on that field to the retriever, so each photo comes back once, at the rank of its best crop. Prices are on the pricing page.
For other approaches to the same problem, see mask-aware retrieval, which segments objects before searching, and open-vocabulary object detection. To find the exact same item rather than similar ones, see instance-level visual matching. To pick the embedding model, see the best multimodal embedding models and the best AI image search tools.
Frequently Asked Questions
Why does my image search miss small objects in a photo?
It stores one vector per photo, computed from a copy shrunk to about 224 pixels. A small object becomes a few pixels, less than one of the model's patches, and the vector mostly describes the main subject. Embedding crops of the photo, and ranking each photo by its best crop, lets the small object match.
Why does CLIP ignore objects at the edge of an image?
CLIP's preprocessing resizes the short side to 224 pixels and then crops the center 224x224 square. On a 4:3 photo that removes a quarter of the width, so anything near the left or right edge never reaches the model.
Should I tile my images or run an object detector first?
Tile when you can't predict what people will search for, since tiling covers the whole photo. Detect first when the objects come from a known set, such as products or logos, because you store far fewer vectors. Either way, keep a whole-photo vector for scene-level queries.
Does a bigger or higher-resolution embedding model solve it?
Partly. SigLIP 2 NaFlex keeps the aspect ratio and can read more patches, so less detail is lost. It still makes one vector for the whole photo, and the rest of the scene outweighs a small object in it.