How Hierarchical Index Navigation Works in kb-retriever: A Deep Dive

The kb-retriever skill navigates large knowledge bases by recursively traversing data_structure.md index files at each directory level, scoring child entries against the query, and descending only into the most relevant branches to minimize token usage.

The kb-retriever skill in the ConardLi/garden-skills repository implements an intelligent document retrieval system that avoids loading entire knowledge bases into context. Instead, it treats the file system as a hierarchical index where each directory contains a structured map of its contents.

The Hierarchical Index Architecture

The Role of data_structure.md Files

At the heart of the navigation system lies the data_structure.md file. Every directory in the knowledge base contains one of these index files, which acts as a local table of contents. According to the skill's documentation in skills/kb-retriever/README.md, these files enumerate sub-directories and files along with metadata describing their purpose and coverage.

The skill reads these index files to understand the semantic structure of the knowledge base without opening the actual documents. This design pattern allows the agent to prune irrelevant branches early in the retrieval process, preventing the "fan-out" problem common in naive recursive searches.

Directory Tree as Navigation Graph

The file system transforms into a directed graph where:

  • Nodes represent directories containing data_structure.md files
  • Edges represent the parent-child relationships described within those index files
  • Leaf nodes are actual knowledge files (Markdown, PDF, Excel, or text files)

As noted in the source comments at lines 94-100 of skills/kb-retriever/README.md, the skill "walks a hierarchical index of data_structure.md files to figure out which files are likely to contain the answer."

How the Navigation Algorithm Works

Step 1: Root Index Evaluation

The process begins at the root directory of the knowledge base. The skill locates and parses the root data_structure.md to extract all immediate children (sub-directories and files).

def evaluate_root_index(kb_path: str, query: str) -> list:
    """Read the root index and identify candidate branches."""
    index_file = os.path.join(kb_path, "data_structure.md")
    
    if not os.path.isfile(index_file):
        return []  # No index available; cannot navigate

    
    entries = parse_structure_md(index_file)  # Extract headings and metadata

    return score_entries(entries, query)

Step 2: Relevance Scoring and Pruning

Each entry in the index receives a relevance score based on keyword overlap between the entry description and the user query. The skill considers metadata fields such as coverage areas, document purpose, and semantic tags defined in the data_structure.md format.

Only the top-scoring candidates (typically 1-2 entries) are retained for further exploration. This aggressive pruning ensures the recursion stays focused and prevents the combinatorial explosion of traversing every branch.

Step 3: Recursive Descent

For each selected candidate that is a directory, the skill recursively descends and repeats the evaluation process:

def walk_index(current_path: str, query: str, depth: int = 0, max_depth: int = 5) -> list:
    """Recursively navigate the hierarchical index."""
    index_path = os.path.join(current_path, "data_structure.md")
    
    if not os.path.exists(index_path) or depth >= max_depth:
        return [current_path]  # Return leaf for content retrieval

    
    entries = parse_index(index_path)
    scored = [(e, calculate_relevance(e, query)) for e in entries]
    scored.sort(key=lambda x: x[1], reverse=True)
    
    # Limit branching factor to prevent fan-out

    top_candidates = [e for e, _ in scored[:2]]
    
    results = []
    for candidate in top_candidates:
        candidate_path = os.path.join(current_path, candidate.name)
        if candidate.is_directory:
            results.extend(walk_index(candidate_path, query, depth + 1, max_depth))
        else:
            results.append(candidate_path)
    
    return results

Step 4: Leaf Node Retrieval

When the algorithm reaches a leaf file (a non-directory entry) or a directory without a data_structure.md, it terminates the navigation phase. The skill then performs progressive retrieval on the selected file(s), using grep-first strategies and windowed reads to extract only the relevant sections, as detailed in the skills/kb-retriever/references/pdf_reading.md guidance.

Implementation Details from the Source

The core logic is documented in skills/kb-retriever/README.md at lines 94-100, which explains: "For each directory level the skill reads data_structure.md, picks the most relevant child(ren) for the user's question, and recurses — so it doesn't fan out across the whole tree."

This implementation choice reflects a best-first search strategy rather than breadth-first or depth-first traversal, optimizing for token efficiency in large knowledge bases. The skill also includes fallback mechanisms for handling scanned PDFs via skills/kb-retriever/scripts/convert_pdf_to_images.py when text extraction fails.

File Structure and Key Components

Summary

  • Hierarchical index navigation in kb-retriever uses data_structure.md files as semantic signposts at every directory level
  • The algorithm performs relevance scoring at each node and descends only into the highest-scoring branches, preventing token waste
  • Recursive traversal continues until reaching leaf files, at which point the skill applies format-specific progressive retrieval
  • This architecture supports multi-format knowledge bases including Markdown, PDF, and Excel without loading entire documents into context

Frequently Asked Questions

How does kb-retriever avoid searching the entire knowledge base?

The skill implements aggressive pruning at each directory level. By reading the local data_structure.md index and scoring only immediate children, it eliminates irrelevant branches before descending. As implemented in ConardLi/garden-skills, this ensures only 1-2 promising paths are explored rather than the entire tree.

What format must the index files follow?

Each directory must contain a data_structure.md file that lists sub-directories and files with descriptive metadata. The skill parses these using keyword matching against the user query to determine relevance scores for navigation decisions.

Does kb-retriever work with binary documents like PDFs and Excel files?

Yes. The navigation algorithm treats these as leaf nodes in the hierarchical index. Once the algorithm selects a PDF or Excel file via the index traversal, it delegates to specialized handlers defined in skills/kb-retriever/references/pdf_reading.md and excel_reading.md, which use tools like pdftotext or pdfplumber for progressive content extraction without loading the entire file.

What happens if a directory lacks a data_structure.md file?

The skill treats directories without an index file as termination points. It either returns the directory path for manual inspection or attempts direct file listing, depending on the configuration. However, the hierarchical navigation pattern requires these index files for optimal performance, as noted in the skill's implementation guidelines.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →