Pick the metric that matches user behavior
Pick the metric that matches user behavior

How do you know your search works? Most teams answer with a shrug and a demo query.
Retrieval has real metrics, and picking one is a product decision, not a math one. If users trust the first result (support bots, agents), measure MRR: how high the first correct answer ranks. If they scan a page of results, NDCG@10: whether the good stuff sits near the top. If missing anything is expensive (legal discovery, brand safety), recall@k: what fraction of the true matches came back.
Optimize the wrong one and things get worse invisibly. Chasing recall floods agents with context they never read. Chasing MRR hides the third relevant document a lawyer needed.
We built eval directly into Mixpeek retrievers because the alternative is what most pipelines do today: ship, vibe-check, drift, rot. If your eval can fail silently, you don't have quality. You have luck.
Unmeasured retrieval isn't working. It's just not failing loudly yet.
Where this diagram appears
Run this on your own data
Mixpeek turns video, images, audio, and documents in your object storage into searchable, timestamped results through one API.
Search your own data, free

