NEWVectors or files. Pick a path.Start →
    Retrieval
    8 min read
    Updated 2026-09-28

    Which Retrieval Pipeline Steps Should Use a Decision Model Like Jev?

    Many steps in a retrieval pipeline are decisions over a fixed set of options: keep or discard, pick a label, rate one to five, route to a branch. A decision model such as TypeSafe's Jev answers these as calibrated probabilities, in one fast call, at a fraction of the cost of generating text. This explains which steps fit, how to phrase them, and where a generative model is still the right tool.

    Jev
    Classification
    Reranking
    LLM Filtering
    Cluster Labeling
    Retrieval

    The Short Answer



    Use a decision model for any pipeline step whose output is a choice from a set you can list: keep or discard a document, pick a category, rate relevance on a scale, choose which branch of a hierarchy to search. Summaries, fresh cluster names and answers are new text, and they stay with a generative LLM.

    A decision model such as TypeSafe's Jev takes the text and a typed question and returns a probability for each option. That makes it cheaper and more predictable for the decision steps, and the probability is useful on its own: it tells you which answers to trust and which to send for review.

    Spotting a decision step



    Look at the output schema of each step. If each field is a boolean, an enum, a number from 0 to 1, or an integer with a small range, the step is a decision.

    StepOutputDecision model fits
    Relevance filterkeep: booleanYes
    Classification at ingestioncategory: enum, is_ad: booleanYes
    Rerankingrelevance probability per documentYes
    Cluster labeling from a vocabularylabel: enumYes
    Routing a query through a cluster treebranch: enum at each levelYes
    Cluster naming from scratchlabel: free textNo
    Summaries, answers, descriptionsfree textNo

    Use the probability



    A decision model answers keep with a number such as 0.98 or 0.51. That number lets you set a threshold per use case, send the uncertain middle to a person or to a larger model, and notice when a batch of documents is more ambiguous than usual. In a reranker the probability is the score itself.

    In our own tests, a filter for "product complaints" kept a complaint at 0.98 and dropped a recipe at 0.01. In production, a reranker asked for "food or a food gift" put four food products at the top of a 30-document candidate list, each at 0.98, in 377 ms, and a classifier over eight product categories got six of six right.

    Phrasing the question



    Decision models read the question word for word. Write the exact condition:

  1. "The document states the refund policy for damaged items" works. "Relevant to
  2. refunds" produces vague scores.
  3. Keep arithmetic, counting and date comparison in code. Ask the model for the parts
  4. and compute the answer yourself.
  5. Send only the text the question needs. Unrelated text lowers accuracy.


  6. Cost



    TypeSafe prices Jev at $0.042 per million input tokens, with output free. A filter decision over a short document is a few hundred tokens, so a million decisions costs a few dollars. Jev reads text only, so for video, images and audio it judges the transcript, captions and metadata.

    In Mixpeek



    Mixpeek offers Jev at five decision points: llm_provider: "typesafe" in the text extractor, model_name: "jev-latest" in the LLM Filter and LLM Enrich stages, inference_name: "typesafe__jev" in the Rerank stage, fixed-vocabulary cluster labeling through candidate_labels, and the cluster_navigation strategy for walking a cluster hierarchy.

    Related



  7. How do you run agentic search over a hierarchy of clusters?
  8. Jev in Mixpeek
  9. Managed Mixpeek

    Put multimodal search to work

    Connect a bucket and Mixpeek runs the whole multimodal search pipeline for you: extraction, indexing, and search over your own objects. No models to wire up, nothing to host.

    Start with Managed
    MVS · bring your own

    Already have vectors?

    Keep your embeddings on your own cloud and run dense, sparse, and BM25 search directly on object storage. From $25/mo.

    Start with MVS

    Run this on your own data

    Point Mixpeek at the storage you already have and search your video, images, audio, and documents the way this guide describes. Build starts at $25/mo for up to 1M vectors.

    Search your own archiveRead Docs