> ## Documentation Index
> Fetch the complete documentation index at: https://docs.mixpeek.com/docs/llms.txt
> Use this file to discover all available pages before exploring further.

# Hierarchical search over clusters

> Cluster a collection into a labeled hierarchy, then answer queries by walking it with Jev

Vector search finds the documents nearest to a query. Hierarchical search first finds the right part of your corpus, then searches inside it. This guide clusters a collection into a two-level hierarchy, labels every cluster from a fixed vocabulary, and answers queries by walking the tree top-down with [Jev](/docs/processing/jev).

The walk uses the `cluster_navigation` strategy of [Agent Search](/docs/retrieval/stages/agent-search). At each level, Jev returns a probability for every child cluster. The stage keeps the best few paths and returns members of the best leaf clusters.

## 1. Cluster into a hierarchy and label it

Set `hierarchical: true` to split each top-level cluster into sub-clusters. Give `candidate_labels` the names you want at both levels.

<CodeGroup>
  ```bash cURL theme={null}
  curl -sS -X POST "$MP_API_URL/v1/clusters" \
    -H "Authorization: Bearer $MP_API_KEY" \
    -H "X-Namespace: $MP_NAMESPACE" \
    -H "Content-Type: application/json" \
    -d '{
      "cluster_name": "content_tree",
      "collection_ids": ["col_content"],
      "cluster_type": "vector",
      "vector_config": {
        "feature_uri": "mixpeek://text_extractor@v1/multilingual_e5_large_instruct_v1",
        "clustering_method": "kmeans",
        "algorithm_params": {"n_clusters": 3},
        "hierarchical": true,
        "max_hierarchy_depth": 2
      },
      "llm_labeling": {
        "enabled": true,
        "provider": "typesafe",
        "model_name": "jev-latest",
        "candidate_labels": ["Product Ads", "Sports", "Cooking",
                             "Pharmaceutical Ads", "Car Ads",
                             "Skateboarding", "Surfing", "Baking", "Grilling"],
        "labeling_inputs": {
          "input_mappings": [{"input_key": "text", "source_type": "payload", "path": "content"}]
        }
      }
    }'

  curl -sS -X POST "$MP_API_URL/v1/clusters/$CLUSTER_ID/execute" \
    -H "Authorization: Bearer $MP_API_KEY" \
    -H "X-Namespace: $MP_NAMESPACE"
  ```

  ```python Python theme={null}
  import requests

  H = {"Authorization": f"Bearer {MP_API_KEY}", "X-Namespace": MP_NAMESPACE}
  cluster = requests.post(f"{MP_API_URL}/v1/clusters", headers=H, json={
      "cluster_name": "content_tree",
      "collection_ids": ["col_content"],
      "cluster_type": "vector",
      "vector_config": {
          "feature_uri": "mixpeek://text_extractor@v1/multilingual_e5_large_instruct_v1",
          "clustering_method": "kmeans",
          "algorithm_params": {"n_clusters": 3},
          "hierarchical": True,
          "max_hierarchy_depth": 2,
      },
      "llm_labeling": {
          "enabled": True,
          "provider": "typesafe",
          "model_name": "jev-latest",
          "candidate_labels": ["Product Ads", "Sports", "Cooking",
                               "Pharmaceutical Ads", "Car Ads",
                               "Skateboarding", "Surfing", "Baking", "Grilling"],
          "labeling_inputs": {"input_mappings": [
              {"input_key": "text", "source_type": "payload", "path": "content"}]},
      },
  }).json()
  requests.post(f"{MP_API_URL}/v1/clusters/{cluster['cluster_id']}/execute", headers=H)
  ```

  ```javascript JavaScript theme={null}
  const H = {
    Authorization: `Bearer ${MP_API_KEY}`,
    "X-Namespace": MP_NAMESPACE,
    "Content-Type": "application/json",
  };
  const cluster = await (await fetch(`${MP_API_URL}/v1/clusters`, {
    method: "POST",
    headers: H,
    body: JSON.stringify({
      cluster_name: "content_tree",
      collection_ids: ["col_content"],
      cluster_type: "vector",
      vector_config: {
        feature_uri: "mixpeek://text_extractor@v1/multilingual_e5_large_instruct_v1",
        clustering_method: "kmeans",
        algorithm_params: { n_clusters: 3 },
        hierarchical: true,
        max_hierarchy_depth: 2,
      },
      llm_labeling: {
        enabled: true,
        provider: "typesafe",
        model_name: "jev-latest",
        candidate_labels: ["Product Ads", "Sports", "Cooking", "Pharmaceutical Ads",
          "Car Ads", "Skateboarding", "Surfing", "Baking", "Grilling"],
        labeling_inputs: { input_mappings: [
          { input_key: "text", source_type: "payload", path: "content" }] },
      },
    }),
  })).json();
  await fetch(`${MP_API_URL}/v1/clusters/${cluster.cluster_id}/execute`, { method: "POST", headers: H });
  ```
</CodeGroup>

The run writes one centroid document per cluster into the cluster's output collection. Each centroid carries `label`, `label_confidence` and `parent_cluster_id`.

## 2. Create a retriever that walks the tree

Point the retriever at the cluster's output collection.

<CodeGroup>
  ```bash cURL theme={null}
  curl -sS -X POST "$MP_API_URL/v1/retrievers" \
    -H "Authorization: Bearer $MP_API_KEY" \
    -H "X-Namespace: $MP_NAMESPACE" \
    -H "Content-Type: application/json" \
    -d '{
      "retriever_name": "browse_content_tree",
      "collection_identifiers": ["'"$OUTPUT_COLLECTION_ID"'"],
      "input_schema": {"query": {"type": "text", "required": true}},
      "stages": [{
        "stage_name": "browse",
        "config": {
          "stage_id": "agent_search",
          "parameters": {
            "strategy": "cluster_navigation",
            "model_name": "jev-latest",
            "cluster_navigation": {"beam_width": 2, "max_leaf_clusters": 1}
          }
        }
      }]
    }'
  ```

  ```python Python theme={null}
  retriever = requests.post(f"{MP_API_URL}/v1/retrievers", headers=H, json={
      "retriever_name": "browse_content_tree",
      "collection_identifiers": [OUTPUT_COLLECTION_ID],
      "input_schema": {"query": {"type": "text", "required": True}},
      "stages": [{"stage_name": "browse", "config": {
          "stage_id": "agent_search",
          "parameters": {
              "strategy": "cluster_navigation",
              "model_name": "jev-latest",
              "cluster_navigation": {"beam_width": 2, "max_leaf_clusters": 1},
          },
      }}],
  }).json()
  ```

  ```javascript JavaScript theme={null}
  const retriever = await (await fetch(`${MP_API_URL}/v1/retrievers`, {
    method: "POST",
    headers: H,
    body: JSON.stringify({
      retriever_name: "browse_content_tree",
      collection_identifiers: [OUTPUT_COLLECTION_ID],
      input_schema: { query: { type: "text", required: true } },
      stages: [{ stage_name: "browse", config: {
        stage_id: "agent_search",
        parameters: {
          strategy: "cluster_navigation",
          model_name: "jev-latest",
          cluster_navigation: { beam_width: 2, max_leaf_clusters: 1 },
        },
      } }],
    }),
  })).json();
  ```
</CodeGroup>

## 3. Query it

<CodeGroup>
  ```bash cURL theme={null}
  curl -sS -X POST "$MP_API_URL/v1/retrievers/$RETRIEVER_ID/execute" \
    -H "Authorization: Bearer $MP_API_KEY" \
    -H "X-Namespace: $MP_NAMESPACE" \
    -H "Content-Type: application/json" \
    -d '{"inputs": {"query": "drug commercials that disclose side effects"}}'
  ```

  ```python Python theme={null}
  result = requests.post(
      f"{MP_API_URL}/v1/retrievers/{retriever['retriever_id']}/execute",
      headers=H,
      json={"inputs": {"query": "drug commercials that disclose side effects"}},
  ).json()
  ```

  ```javascript JavaScript theme={null}
  const result = await (await fetch(
    `${MP_API_URL}/v1/retrievers/${retriever.retriever_id}/execute`,
    { method: "POST", headers: H,
      body: JSON.stringify({ inputs: { query: "drug commercials that disclose side effects" } }) },
  )).json();
  ```
</CodeGroup>

The stage metadata shows the path Jev took. `chosen_clusters` lists each leaf with its path and score, and `navigation_trace` lists the candidates and probabilities at every level.

```json theme={null}
{
  "chosen_clusters": [
    {"cluster_id": "cl_3", "label": "Gourmet Food & Floral Arrangements",
     "path": ["Gourmet Food & Floral Arrangements"], "score": 1.0}
  ]
}
```

That output is from a production run of "a hand-tied bouquet of flowers" against a
product catalog: the stage returned 10 members, all food or home products, in 415 ms.

## Combine with vector search

Put a `feature_search` stage before `agent_search` to search first and then keep only the hits inside the chosen clusters. The order from the search is preserved. Member documents need a cluster field, so enrich the source collection with cluster assignments (`enrich_source_collection: true`) and set `cluster_navigation.centroid_collection_ids` to the collection holding the centroids.

## Tuning

* `beam_width: 1` is greedy descent and makes one request per level. Wider beams let a later level correct an ambiguous early choice. All beams at one level share one request.
* Summaries improve routing. Jev sees each centroid's `summary` as the option's description.
* `max_leaf_clusters` above 1 returns neighbors of the best cluster, useful when a query spans topics.
