How to Answer Questions from a Local Knowledge Base Using kb-retriever
The kb-retriever skill employs hierarchical index navigation and progressive grep-based retrieval to answer natural-language queries over local multi-format documents without loading entire files into the LLM context window.
The kb-retriever skill in the ConardLi/garden-skills repository provides a token-efficient engine to answer questions from a local knowledge base using kb-retriever while maintaining strict context window limits. By traversing directory-specific data_structure.md indices and retrieving only relevant line windows, this tool enables AI agents to reason over large corpora of PDFs, Excel files, and markdown documents with full source attribution.
Architecture of the kb-retriever Skill
The skill operates through four distinct logical layers that work together to minimize token consumption while maximizing retrieval accuracy.
Knowledge Base Layout and Indexing
At the root of the system lies the knowledge/ folder (or any user-specified path) containing a hierarchy of sub-folders. Each directory maintains a data_structure.md index file that describes its purpose, contained files, and coverage scope. According to the source code in skills/kb-retriever/README.md (lines 46-65), these index files form a traversable tree that the skill navigates to narrow down candidate files without scanning entire directory listings.
Hierarchical Index Navigation
When a query arrives, the skill reads the data_structure.md of the current directory and selects the most relevant child directories or files based on semantic relevance to the query. It then recurses deeper into the selected branches. As documented in the README (lines 94-100), this index-walking approach ensures token usage remains low even for large corpora because the skill never loads file contents until the final retrieval stage.
The Four-Stage Retrieval Pipeline
1. Directory Tree Traversal
The skill begins by locating the root knowledge folder (default knowledge/ or custom path specified via the path parameter). It parses the data_structure.md at each level to determine which subdirectories likely contain answers, effectively pruning irrelevant branches of the document tree before any content extraction occurs.
2. Progressive Content Retrieval
For each candidate file identified during traversal, the skill executes a lightweight grep-style search to locate specific line windows containing query terms. Only these targeted windows—specified by offset and limit parameters—are read and fed to the LLM. As implemented in the source (README.md lines 13-16), this design guarantees that entire documents never enter the prompt, regardless of file size.
3. Learn-Before-Process for Complex Formats
When encountering binary formats like PDF or Excel, the skill first loads the appropriate reference documentation:
references/pdf_reading.mdfor PDF extraction strategiesreferences/excel_reading.mdfor pandas-based table reading
These reference files explain how to extract text or tabular data safely using tools like pdftotext or pandas. Only after this "learning" step does the skill invoke the appropriate tool to obtain manageable text fragments (README.md lines 13-15).
4. Bounded Iterative Refinement
The retrieval loop executes a maximum of five rounds, with each round refining the keyword set based on intermediate findings. This bounded approach ensures the process terminates predictably while gathering sufficient evidence to produce a final answer with explicit source citations to specific files and line ranges.
Installing and Running kb-retriever
Install the skill using the npm-based skills CLI:
npx skills add ConardLi/garden-skills -s kb-retriever -a claude-code
Querying Excel Files
To answer questions from a local knowledge base using kb-retriever against spreadsheet data:
echo '{"query":"What are the quarterly sales trends for 2023?","path":"./docs"}' \
| npx skills run kb-retriever
The skill executes the following sequence:
- Detects the
./docsfolder (falling back toknowledge/if unspecified) - Walks the
data_structure.mdhierarchy inside that folder - Performs grep searches for "quarterly sales" in candidate Excel files
- Reads only relevant rows (e.g., rows 15-27) and returns an answer like:
The Q1-Q4 2023 sales increased by 12% YoY (source:
sales_2023.xlsx– rows 15-27).
Querying PDF Documents
For PDF-specific queries:
echo '{"query":"Summarize the legal compliance checklist in the policy.pdf"}' \
| npx skills run kb-retriever
The skill first consults references/pdf_reading.md to determine whether to use pdftotext or pdfplumber, then extracts only the relevant sections and answers with a citation to policy.pdf.
Key Source Files and References
The following files implement the kb-retriever workflow in the ConardLi/garden-skills repository:
skills/kb-retriever/README.md– Complete user-facing documentation including architecture overview and usage instructions (lines 46-65, 94-100).skills/kb-retriever/SKILL.md– Skill manifest containing frontmatter name and description.skills/kb-retriever/references/pdf_reading.md– Guidance on PDF text extraction, tool selection, and fallback to image conversion.skills/kb-retriever/references/excel_reading.md– Instructions for safely reading Excel files with pandas, including row limits and dtype handling.skills/kb-retriever/scripts/convert_pdf_to_images.py– Helper script that converts PDF pages to images when text extraction fails, enabling visual processing as a fallback mechanism.
Summary
- Hierarchical indexing via
data_structure.mdfiles enables scalable navigation without loading full directory listings or file contents. - Progressive retrieval uses grep-style searches to extract only relevant line windows, ensuring large documents never consume excessive tokens.
- Learn-before-process pattern handles complex binary formats by first reading reference documentation to select appropriate extraction tools.
- Five-round retrieval limit provides bounded execution with iterative keyword refinement and explicit source citations for traceability.
Frequently Asked Questions
What file formats does kb-retriever support?
The skill supports plain text, markdown, PDF, and Excel files. For binary formats like PDF and Excel, it consults references/pdf_reading.md and references/excel_reading.md to determine the appropriate extraction tool (such as pdftotext, pdfplumber, or pandas) before attempting to read content.
How does kb-retriever avoid exceeding LLM token limits?
Instead of loading entire documents into the context window, the skill traverses hierarchical data_structure.md indices to locate candidate files, then performs targeted grep searches to retrieve only small line windows containing query terms. This progressive retrieval strategy ensures that only highly relevant text fragments enter the LLM prompt.
Can I use a custom knowledge base folder instead of the default knowledge/ directory?
Yes. Specify the path property in your JSON input when invoking the skill. The tool will detect your custom folder and traverse its internal data_structure.md hierarchy accordingly, treating it exactly like the default knowledge/ root.
What happens when text extraction fails on a PDF document?
The skill references skills/kb-retriever/scripts/convert_pdf_to_images.py as a fallback mechanism. If primary text extraction tools fail, the agent can convert PDF pages to images for visual processing, following the guidance provided in references/pdf_reading.md for image-based content extraction.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →