# How Hierarchical Index Navigation Works in kb-retriever: A Deep Dive

> Discover how kb-retriever navigates vast knowledge bases using hierarchical index navigation. Learn to traverse, score, and descend into relevant branches efficiently to save tokens.

- Repository: [ConardLi/garden-skills](https://github.com/ConardLi/garden-skills)
- Tags: deep-dive
- Published: 2026-08-31

---

**The kb-retriever skill navigates large knowledge bases by recursively traversing [`data_structure.md`](https://github.com/ConardLi/garden-skills/blob/main/data_structure.md) index files at each directory level, scoring child entries against the query, and descending only into the most relevant branches to minimize token usage.**

The **kb-retriever** skill in the `ConardLi/garden-skills` repository implements an intelligent document retrieval system that avoids loading entire knowledge bases into context. Instead, it treats the file system as a **hierarchical index** where each directory contains a structured map of its contents.

## The Hierarchical Index Architecture

### The Role of data_structure.md Files

At the heart of the navigation system lies the **[`data_structure.md`](https://github.com/ConardLi/garden-skills/blob/main/data_structure.md)** file. Every directory in the knowledge base contains one of these index files, which acts as a local table of contents. According to the skill's documentation in [`skills/kb-retriever/README.md`](https://github.com/ConardLi/garden-skills/blob/main/skills/kb-retriever/README.md), these files enumerate sub-directories and files along with metadata describing their purpose and coverage.

The skill reads these index files to understand the semantic structure of the knowledge base without opening the actual documents. This design pattern allows the agent to **prune irrelevant branches** early in the retrieval process, preventing the "fan-out" problem common in naive recursive searches.

### Directory Tree as Navigation Graph

The file system transforms into a directed graph where:
- **Nodes** represent directories containing [`data_structure.md`](https://github.com/ConardLi/garden-skills/blob/main/data_structure.md) files
- **Edges** represent the parent-child relationships described within those index files
- **Leaf nodes** are actual knowledge files (Markdown, PDF, Excel, or text files)

As noted in the source comments at lines 94-100 of [`skills/kb-retriever/README.md`](https://github.com/ConardLi/garden-skills/blob/main/skills/kb-retriever/README.md), the skill "walks a hierarchical index of [`data_structure.md`](https://github.com/ConardLi/garden-skills/blob/main/data_structure.md) files to figure out which files are likely to contain the answer."

## How the Navigation Algorithm Works

### Step 1: Root Index Evaluation

The process begins at the **root directory** of the knowledge base. The skill locates and parses the root [`data_structure.md`](https://github.com/ConardLi/garden-skills/blob/main/data_structure.md) to extract all immediate children (sub-directories and files).

```python
def evaluate_root_index(kb_path: str, query: str) -> list:
    """Read the root index and identify candidate branches."""
    index_file = os.path.join(kb_path, "data_structure.md")
    
    if not os.path.isfile(index_file):
        return []  # No index available; cannot navigate

    
    entries = parse_structure_md(index_file)  # Extract headings and metadata

    return score_entries(entries, query)

```

### Step 2: Relevance Scoring and Pruning

Each entry in the index receives a **relevance score** based on keyword overlap between the entry description and the user query. The skill considers metadata fields such as coverage areas, document purpose, and semantic tags defined in the [`data_structure.md`](https://github.com/ConardLi/garden-skills/blob/main/data_structure.md) format.

Only the **top-scoring candidates** (typically 1-2 entries) are retained for further exploration. This aggressive pruning ensures the recursion stays focused and prevents the combinatorial explosion of traversing every branch.

### Step 3: Recursive Descent

For each selected candidate that is a directory, the skill **recursively descends** and repeats the evaluation process:

```python
def walk_index(current_path: str, query: str, depth: int = 0, max_depth: int = 5) -> list:
    """Recursively navigate the hierarchical index."""
    index_path = os.path.join(current_path, "data_structure.md")
    
    if not os.path.exists(index_path) or depth >= max_depth:
        return [current_path]  # Return leaf for content retrieval

    
    entries = parse_index(index_path)
    scored = [(e, calculate_relevance(e, query)) for e in entries]
    scored.sort(key=lambda x: x[1], reverse=True)
    
    # Limit branching factor to prevent fan-out

    top_candidates = [e for e, _ in scored[:2]]
    
    results = []
    for candidate in top_candidates:
        candidate_path = os.path.join(current_path, candidate.name)
        if candidate.is_directory:
            results.extend(walk_index(candidate_path, query, depth + 1, max_depth))
        else:
            results.append(candidate_path)
    
    return results

```

### Step 4: Leaf Node Retrieval

When the algorithm reaches a **leaf file** (a non-directory entry) or a directory without a [`data_structure.md`](https://github.com/ConardLi/garden-skills/blob/main/data_structure.md), it terminates the navigation phase. The skill then performs **progressive retrieval** on the selected file(s), using grep-first strategies and windowed reads to extract only the relevant sections, as detailed in the [`skills/kb-retriever/references/pdf_reading.md`](https://github.com/ConardLi/garden-skills/blob/main/skills/kb-retriever/references/pdf_reading.md) guidance.

## Implementation Details from the Source

The core logic is documented in [`skills/kb-retriever/README.md`](https://github.com/ConardLi/garden-skills/blob/main/skills/kb-retriever/README.md) at lines 94-100, which explains: "For each directory level the skill reads [`data_structure.md`](https://github.com/ConardLi/garden-skills/blob/main/data_structure.md), picks the most relevant child(ren) for the user's question, and recurses — so it doesn't fan out across the whole tree."

This implementation choice reflects a **best-first search** strategy rather than breadth-first or depth-first traversal, optimizing for token efficiency in large knowledge bases. The skill also includes fallback mechanisms for handling scanned PDFs via [`skills/kb-retriever/scripts/convert_pdf_to_images.py`](https://github.com/ConardLi/garden-skills/blob/main/skills/kb-retriever/scripts/convert_pdf_to_images.py) when text extraction fails.

## File Structure and Key Components

- **[`skills/kb-retriever/README.md`](https://github.com/ConardLi/garden-skills/blob/main/skills/kb-retriever/README.md)** — Contains the primary documentation for the hierarchical navigation algorithm (lines 94-100)
- **[`skills/kb-retriever/SKILL.md`](https://github.com/ConardLi/garden-skills/blob/main/skills/kb-retriever/SKILL.md)** — Defines the skill metadata including version 0.2.1 and the description of navigating "a hierarchical index of knowledge files"
- **[`skills/kb-retriever/references/pdf_reading.md`](https://github.com/ConardLi/garden-skills/blob/main/skills/kb-retriever/references/pdf_reading.md)** — Provides tooling guidance for progressive PDF text extraction
- **[`skills/kb-retriever/references/excel_reading.md`](https://github.com/ConardLi/garden-skills/blob/main/skills/kb-retriever/references/excel_reading.md)** — Specifies pandas-based strategies for Excel file retrieval
- **[`data_structure.md`](https://github.com/ConardLi/garden-skills/blob/main/data_structure.md)** (template) — The index file format placed in each knowledge base directory to enable navigation

## Summary

- **Hierarchical index navigation** in kb-retriever uses [`data_structure.md`](https://github.com/ConardLi/garden-skills/blob/main/data_structure.md) files as semantic signposts at every directory level
- The algorithm performs **relevance scoring** at each node and descends only into the highest-scoring branches, preventing token waste
- **Recursive traversal** continues until reaching leaf files, at which point the skill applies format-specific progressive retrieval
- This architecture supports multi-format knowledge bases including Markdown, PDF, and Excel without loading entire documents into context

## Frequently Asked Questions

### How does kb-retriever avoid searching the entire knowledge base?

The skill implements **aggressive pruning** at each directory level. By reading the local [`data_structure.md`](https://github.com/ConardLi/garden-skills/blob/main/data_structure.md) index and scoring only immediate children, it eliminates irrelevant branches before descending. As implemented in `ConardLi/garden-skills`, this ensures only 1-2 promising paths are explored rather than the entire tree.

### What format must the index files follow?

Each directory must contain a **[`data_structure.md`](https://github.com/ConardLi/garden-skills/blob/main/data_structure.md)** file that lists sub-directories and files with descriptive metadata. The skill parses these using keyword matching against the user query to determine relevance scores for navigation decisions.

### Does kb-retriever work with binary documents like PDFs and Excel files?

Yes. The navigation algorithm treats these as **leaf nodes** in the hierarchical index. Once the algorithm selects a PDF or Excel file via the index traversal, it delegates to specialized handlers defined in [`skills/kb-retriever/references/pdf_reading.md`](https://github.com/ConardLi/garden-skills/blob/main/skills/kb-retriever/references/pdf_reading.md) and [`excel_reading.md`](https://github.com/ConardLi/garden-skills/blob/main/excel_reading.md), which use tools like `pdftotext` or `pdfplumber` for progressive content extraction without loading the entire file.

### What happens if a directory lacks a data_structure.md file?

The skill treats directories without an index file as **termination points**. It either returns the directory path for manual inspection or attempts direct file listing, depending on the configuration. However, the hierarchical navigation pattern requires these index files for optimal performance, as noted in the skill's implementation guidelines.