NEWVectors or files. Pick a path.Start →
    Retrieval
    9 min read
    Updated 2026-09-28

    How Do You Run Agentic Search Over a Hierarchy of Clusters?

    Cluster the corpus into a labeled tree, then answer a query by walking the tree from the top, choosing the child cluster most likely to hold the answer at every level. A decision model such as TypeSafe's Jev returns a probability for each branch, so the walk can keep several paths open and rank them. This covers how to build the tree, how the walk works, and when it beats a single vector search.

    Clustering
    Hierarchical Search
    Agentic Search
    Jev
    Beam Search
    Retrieval

    The Short Answer



    Cluster your collection into a hierarchy, give each cluster a label, and search by walking that hierarchy from the top. At each level a model reads the query and the labels of the child clusters and picks the one most likely to contain the answer. At the bottom level, you search or return the documents inside the clusters it chose.

    The model making those choices picks one option from a short list and reports how sure it is. An evaluation model such as TypeSafe's Jev is built for that job: each branch decision comes back as a probability for each child cluster. With probabilities you can keep the best two or three paths open at each level and rank complete paths at the end, which is a beam search over your corpus.

    In Mixpeek this is the cluster_navigation strategy of the agent_search retriever stage, with model_name: "jev-latest".

    Cases a single vector search misses



    A single vector search compares the query against the indexed documents and returns the nearest ones. That works when the words in the query sit close to the words in the answer. It struggles in three situations:

  1. The query names a category, and the documents use other words. "Drug
  2. commercials that disclose side effects" shares almost no words with an ad that says "ask your doctor, may cause dizziness", and both belong to the same category.
  3. Neighbouring topics share vocabulary. A sneaker ad featuring a basketball star
  4. and a clip of a basketball game use the same words. One belongs under advertising, the other under sports.
  5. You need to explain the result. A path such as Product Ads > Pharmaceutical Ads
  6. tells a reviewer why a document came back.

    Walking a labeled tree decides the category first and searches second.

    Step 1: build and label the tree



    Run hierarchical clustering over the collection's embeddings. KMeans or HDBSCAN at the top level, then the same algorithm inside each top-level cluster, gives you a two-level tree. Each cluster gets a centroid document that records its label, summary and parent.

    Label the clusters from a vocabulary you control. Generated labels drift between runs and between levels. A fixed list keeps "Pharmaceutical Ads" spelled the same way on each run, and it lets you line the tree up with a taxonomy you already use. Jev picks each cluster's label from the list and returns the probability, so a cluster labeled with low confidence is easy to find and review.

    Step 2: walk it with probabilities



    At the top level the model sees the query and the top-level labels, with each cluster's summary as the description of that option. It returns a probability for each. Keep the best few. At the next level, ask about the children of each path you kept, all in one request. Score each complete path by the geometric mean of the probabilities along it, so a shallow leaf and a deep leaf compare on equal terms.

    Greedy descent, keeping one path, is a beam width of one. It makes one request per level and cannot recover from an early mistake. A beam of two or three lets a later level correct an ambiguous first choice at almost no extra latency, because the paths share a request.

    On an eleven-cluster, two-level tree we tested, the walk routed "Nike commercial with a basketball player" to Product Ads > Sneaker Ads at 0.99, and kept Sports > Basketball as the runner-up at 0.14. Each two-level walk took 0.8 to 1.3 seconds.

    Step 3: return the documents



    Return the members of the best leaf clusters, or run a vector search first and keep only the hits that fall inside the chosen clusters. The second form keeps the vector search's ordering and uses the tree as a filter.

    Use a flat search for short, specific queries



    Use a single vector search when queries name the thing itself ("the Q3 earnings call"), when the corpus is small enough that a query reads most of it anyway, or when latency under 100 ms matters more than precision. Tree walking earns its extra second when queries name categories, when topics overlap, or when the path is part of the answer.

    Related



  7. Mixpeek docs: hierarchical search over clusters
  8. Jev in Mixpeek
  9. Managed Mixpeek

    Put multimodal search to work

    Connect a bucket and Mixpeek runs the whole multimodal search pipeline for you: extraction, indexing, and search over your own objects. No models to wire up, nothing to host.

    Start with Managed
    MVS · bring your own

    Already have vectors?

    Keep your embeddings on your own cloud and run dense, sparse, and BM25 search directly on object storage. From $25/mo.

    Start with MVS

    Run this on your own data

    Point Mixpeek at the storage you already have and search your video, images, audio, and documents the way this guide describes. Build starts at $25/mo for up to 1M vectors.

    Search your own archiveRead Docs

    Related guides

    Retrieval

    Which Retrieval Pipeline Steps Should Use a Decision Model Like Jev?

    Many steps in a retrieval pipeline are decisions over a fixed set of options: keep or discard, pick a label, rate one to five, route to a branch. A decision model such as TypeSafe's Jev answers these as calibrated probabilities, in one fast call, at a fraction of the cost of generating text. This explains which steps fit, how to phrase them, and where a generative model is still the right tool.

    Read guide →
    Retrieval

    Why Do I Get Different Search Results Every Time I Run the Same Query?

    Almost every vector search engine is approximate: it walks a fraction of the index instead of comparing your query against every vector, and the fraction it walks can differ between runs. Five things produce run-to-run variation, they leave different fingerprints, and telling them apart takes about ten minutes. This is how to find which one you have and what to change.

    Read guide →
    Retrieval

    BM25 and the Inverted Index: The Lexical Retriever Every Hybrid Search Treats as a Black Box

    Every hybrid search pipeline pairs dense vectors with BM25, but almost no one can say where the BM25 number actually comes from, which is exactly why fusion, tuning, and exact-match failures stay mysterious. This guide opens the box: how an inverted index turns transcripts and OCR text into posting lists, the precise BM25 scoring formula with its term-frequency saturation and length normalization, what the k1 and b parameters really do, and why the tokenizer is the silent decider of whether an agent ever finds a serial number.

    Read guide →