Skip to main content
The Score Threshold stage applies an absolute quality gate: it drops every document whose score on a chosen field fails a minimum bar. When nothing clears the bar it returns an empty result set and sets all_below_threshold — the signal your UI uses to show a “no good results” state instead of presenting weak matches.
Stage Category: REDUCE (N documents → ≤ N documents)Transformation: keeps only documents meeting min_score; the set may become empty.

Why not just normalize and filter?

score_normalize with min_max always rescales the top result to 1.0 — so a threshold on the normalized score can never reject an all-bad result set (the best match is always 1.0). Score Threshold gates on the raw or calibrated score, so “everything is below the bar” is expressible. Threshold on a calibrated score — the rerank cross-encoder score is ideal (score_field: "scores.rerank"), since it is far more absolute and comparable than a raw cosine similarity.

When to Use

When NOT to Use

Parameters

Response Metadata

Configuration Examples

A raw cosine score does not span 0 to 1, so do not copy a threshold between retrievers. Each embedding model occupies its own narrow band, and a value that works on one model rejects every document on another.Measured on a production multimodal retriever (multimodal_extractor@v2/gemini-embedding-2, video segments, 27 text queries): on-topic queries scored 0.3755 to 0.4311 and off-topic controls scored 0.3136 to 0.3456. The separation is clean, and the whole range sits between 0.31 and 0.44. A min_score of 0.5 or 0.7 on that retriever returns zero documents for every query, including exact matches, and sets all_below_threshold every time.Read your own numbers before you pick a bar. Run the retriever without this stage, look at the top-1 score for a query that should match and one that should not, and put the bar between them. On the retriever above that puts the bar at 0.35, which was verified in both directions: cars driving on a highway keeps all 10 documents with all_below_threshold false, and a library reading room with wooden shelves drops all 10 and sets it true.

”No Good Results → Suggestions” Pattern

Gate on the rerank score; when nothing qualifies, the empty result + all_below_threshold tells your application to fall back to query_expand for adjacent suggestions instead of showing weak matches.
Calibrate min_score empirically: run representative queries (including ones that should return nothing) and pick the value that separates good from bad on the rerank score. The right cutoff is query-domain specific.

Performance

  • Score Normalize - Rescale scores to a common range
  • Rerank - Calibrated cross-encoder scores to gate on
  • Query Expand - Adjacent suggestions when nothing qualifies
  • Limit - Truncate to top-N regardless of score