# How to Answer Questions from a Local Knowledge Base Using kb-retriever

> Learn how to answer questions from a local knowledge base using kb-retriever. This skill efficiently retrieves information from multi-format documents without overwhelming your LLM.

- Repository: [ConardLi/garden-skills](https://github.com/ConardLi/garden-skills)
- Tags: how-to-guide
- Published: 2026-08-31

---

**The kb-retriever skill employs hierarchical index navigation and progressive grep-based retrieval to answer natural-language queries over local multi-format documents without loading entire files into the LLM context window.**

The kb-retriever skill in the ConardLi/garden-skills repository provides a token-efficient engine to answer questions from a local knowledge base using kb-retriever while maintaining strict context window limits. By traversing directory-specific [`data_structure.md`](https://github.com/ConardLi/garden-skills/blob/main/data_structure.md) indices and retrieving only relevant line windows, this tool enables AI agents to reason over large corpora of PDFs, Excel files, and markdown documents with full source attribution.

## Architecture of the kb-retriever Skill

The skill operates through four distinct logical layers that work together to minimize token consumption while maximizing retrieval accuracy.

### Knowledge Base Layout and Indexing

At the root of the system lies the `knowledge/` folder (or any user-specified path) containing a hierarchy of sub-folders. Each directory maintains a [`data_structure.md`](https://github.com/ConardLi/garden-skills/blob/main/data_structure.md) index file that describes its purpose, contained files, and coverage scope. According to the source code in [`skills/kb-retriever/README.md`](https://github.com/ConardLi/garden-skills/blob/main/skills/kb-retriever/README.md) (lines 46-65), these index files form a traversable tree that the skill navigates to narrow down candidate files without scanning entire directory listings.

### Hierarchical Index Navigation

When a query arrives, the skill reads the [`data_structure.md`](https://github.com/ConardLi/garden-skills/blob/main/data_structure.md) of the current directory and selects the most relevant child directories or files based on semantic relevance to the query. It then recurses deeper into the selected branches. As documented in the README (lines 94-100), this index-walking approach ensures token usage remains low even for large corpora because the skill never loads file contents until the final retrieval stage.

## The Four-Stage Retrieval Pipeline

### 1. Directory Tree Traversal

The skill begins by locating the root knowledge folder (default `knowledge/` or custom path specified via the `path` parameter). It parses the [`data_structure.md`](https://github.com/ConardLi/garden-skills/blob/main/data_structure.md) at each level to determine which subdirectories likely contain answers, effectively pruning irrelevant branches of the document tree before any content extraction occurs.

### 2. Progressive Content Retrieval

For each candidate file identified during traversal, the skill executes a lightweight `grep`-style search to locate specific line windows containing query terms. Only these targeted windows—specified by offset and limit parameters—are read and fed to the LLM. As implemented in the source (README.md lines 13-16), this design guarantees that entire documents never enter the prompt, regardless of file size.

### 3. Learn-Before-Process for Complex Formats

When encountering binary formats like PDF or Excel, the skill first loads the appropriate reference documentation:
- [`references/pdf_reading.md`](https://github.com/ConardLi/garden-skills/blob/main/references/pdf_reading.md) for PDF extraction strategies
- [`references/excel_reading.md`](https://github.com/ConardLi/garden-skills/blob/main/references/excel_reading.md) for pandas-based table reading

These reference files explain how to extract text or tabular data safely using tools like `pdftotext` or `pandas`. Only after this "learning" step does the skill invoke the appropriate tool to obtain manageable text fragments (README.md lines 13-15).

### 4. Bounded Iterative Refinement

The retrieval loop executes a maximum of **five rounds**, with each round refining the keyword set based on intermediate findings. This bounded approach ensures the process terminates predictably while gathering sufficient evidence to produce a final answer with explicit source citations to specific files and line ranges.

## Installing and Running kb-retriever

Install the skill using the npm-based skills CLI:

```bash
npx skills add ConardLi/garden-skills -s kb-retriever -a claude-code

```

### Querying Excel Files

To answer questions from a local knowledge base using kb-retriever against spreadsheet data:

```bash
echo '{"query":"What are the quarterly sales trends for 2023?","path":"./docs"}' \
| npx skills run kb-retriever

```

The skill executes the following sequence:
1. Detects the `./docs` folder (falling back to `knowledge/` if unspecified)
2. Walks the [`data_structure.md`](https://github.com/ConardLi/garden-skills/blob/main/data_structure.md) hierarchy inside that folder
3. Performs grep searches for "quarterly sales" in candidate Excel files
4. Reads only relevant rows (e.g., rows 15-27) and returns an answer like:

> The Q1-Q4 2023 sales increased by 12% YoY (source: `sales_2023.xlsx` – rows 15-27).

### Querying PDF Documents

For PDF-specific queries:

```bash
echo '{"query":"Summarize the legal compliance checklist in the policy.pdf"}' \
| npx skills run kb-retriever

```

The skill first consults [`references/pdf_reading.md`](https://github.com/ConardLi/garden-skills/blob/main/references/pdf_reading.md) to determine whether to use `pdftotext` or `pdfplumber`, then extracts only the relevant sections and answers with a citation to `policy.pdf`.

## Key Source Files and References

The following files implement the kb-retriever workflow in the ConardLi/garden-skills repository:

- **[`skills/kb-retriever/README.md`](https://github.com/ConardLi/garden-skills/blob/main/skills/kb-retriever/README.md)** – Complete user-facing documentation including architecture overview and usage instructions (lines 46-65, 94-100).
- **[`skills/kb-retriever/SKILL.md`](https://github.com/ConardLi/garden-skills/blob/main/skills/kb-retriever/SKILL.md)** – Skill manifest containing frontmatter name and description.
- **[`skills/kb-retriever/references/pdf_reading.md`](https://github.com/ConardLi/garden-skills/blob/main/skills/kb-retriever/references/pdf_reading.md)** – Guidance on PDF text extraction, tool selection, and fallback to image conversion.
- **[`skills/kb-retriever/references/excel_reading.md`](https://github.com/ConardLi/garden-skills/blob/main/skills/kb-retriever/references/excel_reading.md)** – Instructions for safely reading Excel files with pandas, including row limits and dtype handling.
- **[`skills/kb-retriever/scripts/convert_pdf_to_images.py`](https://github.com/ConardLi/garden-skills/blob/main/skills/kb-retriever/scripts/convert_pdf_to_images.py)** – Helper script that converts PDF pages to images when text extraction fails, enabling visual processing as a fallback mechanism.

## Summary

- **Hierarchical indexing** via [`data_structure.md`](https://github.com/ConardLi/garden-skills/blob/main/data_structure.md) files enables scalable navigation without loading full directory listings or file contents.
- **Progressive retrieval** uses grep-style searches to extract only relevant line windows, ensuring large documents never consume excessive tokens.
- **Learn-before-process pattern** handles complex binary formats by first reading reference documentation to select appropriate extraction tools.
- **Five-round retrieval limit** provides bounded execution with iterative keyword refinement and explicit source citations for traceability.

## Frequently Asked Questions

### What file formats does kb-retriever support?

The skill supports plain text, markdown, PDF, and Excel files. For binary formats like PDF and Excel, it consults [`references/pdf_reading.md`](https://github.com/ConardLi/garden-skills/blob/main/references/pdf_reading.md) and [`references/excel_reading.md`](https://github.com/ConardLi/garden-skills/blob/main/references/excel_reading.md) to determine the appropriate extraction tool (such as `pdftotext`, `pdfplumber`, or `pandas`) before attempting to read content.

### How does kb-retriever avoid exceeding LLM token limits?

Instead of loading entire documents into the context window, the skill traverses hierarchical [`data_structure.md`](https://github.com/ConardLi/garden-skills/blob/main/data_structure.md) indices to locate candidate files, then performs targeted grep searches to retrieve only small line windows containing query terms. This progressive retrieval strategy ensures that only highly relevant text fragments enter the LLM prompt.

### Can I use a custom knowledge base folder instead of the default `knowledge/` directory?

Yes. Specify the `path` property in your JSON input when invoking the skill. The tool will detect your custom folder and traverse its internal [`data_structure.md`](https://github.com/ConardLi/garden-skills/blob/main/data_structure.md) hierarchy accordingly, treating it exactly like the default `knowledge/` root.

### What happens when text extraction fails on a PDF document?

The skill references [`skills/kb-retriever/scripts/convert_pdf_to_images.py`](https://github.com/ConardLi/garden-skills/blob/main/skills/kb-retriever/scripts/convert_pdf_to_images.py) as a fallback mechanism. If primary text extraction tools fail, the agent can convert PDF pages to images for visual processing, following the guidance provided in [`references/pdf_reading.md`](https://github.com/ConardLi/garden-skills/blob/main/references/pdf_reading.md) for image-based content extraction.