How Hierarchical Index Navigation Works in kb-retriever: A Deep Dive
The kb-retriever skill navigates large knowledge bases by recursively traversing data_structure.md index files at each directory level, scoring child entries against the query, and descending only into the most relevant branches to minimize token usage.
The kb-retriever skill in the ConardLi/garden-skills repository implements an intelligent document retrieval system that avoids loading entire knowledge bases into context. Instead, it treats the file system as a hierarchical index where each directory contains a structured map of its contents.
The Hierarchical Index Architecture
The Role of data_structure.md Files
At the heart of the navigation system lies the data_structure.md file. Every directory in the knowledge base contains one of these index files, which acts as a local table of contents. According to the skill's documentation in skills/kb-retriever/README.md, these files enumerate sub-directories and files along with metadata describing their purpose and coverage.
The skill reads these index files to understand the semantic structure of the knowledge base without opening the actual documents. This design pattern allows the agent to prune irrelevant branches early in the retrieval process, preventing the "fan-out" problem common in naive recursive searches.
Directory Tree as Navigation Graph
The file system transforms into a directed graph where:
- Nodes represent directories containing
data_structure.mdfiles - Edges represent the parent-child relationships described within those index files
- Leaf nodes are actual knowledge files (Markdown, PDF, Excel, or text files)
As noted in the source comments at lines 94-100 of skills/kb-retriever/README.md, the skill "walks a hierarchical index of data_structure.md files to figure out which files are likely to contain the answer."
How the Navigation Algorithm Works
Step 1: Root Index Evaluation
The process begins at the root directory of the knowledge base. The skill locates and parses the root data_structure.md to extract all immediate children (sub-directories and files).
def evaluate_root_index(kb_path: str, query: str) -> list:
"""Read the root index and identify candidate branches."""
index_file = os.path.join(kb_path, "data_structure.md")
if not os.path.isfile(index_file):
return [] # No index available; cannot navigate
entries = parse_structure_md(index_file) # Extract headings and metadata
return score_entries(entries, query)
Step 2: Relevance Scoring and Pruning
Each entry in the index receives a relevance score based on keyword overlap between the entry description and the user query. The skill considers metadata fields such as coverage areas, document purpose, and semantic tags defined in the data_structure.md format.
Only the top-scoring candidates (typically 1-2 entries) are retained for further exploration. This aggressive pruning ensures the recursion stays focused and prevents the combinatorial explosion of traversing every branch.
Step 3: Recursive Descent
For each selected candidate that is a directory, the skill recursively descends and repeats the evaluation process:
def walk_index(current_path: str, query: str, depth: int = 0, max_depth: int = 5) -> list:
"""Recursively navigate the hierarchical index."""
index_path = os.path.join(current_path, "data_structure.md")
if not os.path.exists(index_path) or depth >= max_depth:
return [current_path] # Return leaf for content retrieval
entries = parse_index(index_path)
scored = [(e, calculate_relevance(e, query)) for e in entries]
scored.sort(key=lambda x: x[1], reverse=True)
# Limit branching factor to prevent fan-out
top_candidates = [e for e, _ in scored[:2]]
results = []
for candidate in top_candidates:
candidate_path = os.path.join(current_path, candidate.name)
if candidate.is_directory:
results.extend(walk_index(candidate_path, query, depth + 1, max_depth))
else:
results.append(candidate_path)
return results
Step 4: Leaf Node Retrieval
When the algorithm reaches a leaf file (a non-directory entry) or a directory without a data_structure.md, it terminates the navigation phase. The skill then performs progressive retrieval on the selected file(s), using grep-first strategies and windowed reads to extract only the relevant sections, as detailed in the skills/kb-retriever/references/pdf_reading.md guidance.
Implementation Details from the Source
The core logic is documented in skills/kb-retriever/README.md at lines 94-100, which explains: "For each directory level the skill reads data_structure.md, picks the most relevant child(ren) for the user's question, and recurses — so it doesn't fan out across the whole tree."
This implementation choice reflects a best-first search strategy rather than breadth-first or depth-first traversal, optimizing for token efficiency in large knowledge bases. The skill also includes fallback mechanisms for handling scanned PDFs via skills/kb-retriever/scripts/convert_pdf_to_images.py when text extraction fails.
File Structure and Key Components
skills/kb-retriever/README.md— Contains the primary documentation for the hierarchical navigation algorithm (lines 94-100)skills/kb-retriever/SKILL.md— Defines the skill metadata including version 0.2.1 and the description of navigating "a hierarchical index of knowledge files"skills/kb-retriever/references/pdf_reading.md— Provides tooling guidance for progressive PDF text extractionskills/kb-retriever/references/excel_reading.md— Specifies pandas-based strategies for Excel file retrievaldata_structure.md(template) — The index file format placed in each knowledge base directory to enable navigation
Summary
- Hierarchical index navigation in kb-retriever uses
data_structure.mdfiles as semantic signposts at every directory level - The algorithm performs relevance scoring at each node and descends only into the highest-scoring branches, preventing token waste
- Recursive traversal continues until reaching leaf files, at which point the skill applies format-specific progressive retrieval
- This architecture supports multi-format knowledge bases including Markdown, PDF, and Excel without loading entire documents into context
Frequently Asked Questions
How does kb-retriever avoid searching the entire knowledge base?
The skill implements aggressive pruning at each directory level. By reading the local data_structure.md index and scoring only immediate children, it eliminates irrelevant branches before descending. As implemented in ConardLi/garden-skills, this ensures only 1-2 promising paths are explored rather than the entire tree.
What format must the index files follow?
Each directory must contain a data_structure.md file that lists sub-directories and files with descriptive metadata. The skill parses these using keyword matching against the user query to determine relevance scores for navigation decisions.
Does kb-retriever work with binary documents like PDFs and Excel files?
Yes. The navigation algorithm treats these as leaf nodes in the hierarchical index. Once the algorithm selects a PDF or Excel file via the index traversal, it delegates to specialized handlers defined in skills/kb-retriever/references/pdf_reading.md and excel_reading.md, which use tools like pdftotext or pdfplumber for progressive content extraction without loading the entire file.
What happens if a directory lacks a data_structure.md file?
The skill treats directories without an index file as termination points. It either returns the directory path for manual inspection or attempts direct file listing, depending on the configuration. However, the hierarchical navigation pattern requires these index files for optimal performance, as noted in the skill's implementation guidelines.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →