# How to Use the kb-retriever Skill for Knowledge Retrieval in AI Agents

> Learn to use the kb-retriever skill for efficient knowledge retrieval in AI agents. Explore hierarchical indexes, documentation, and progressive grep for optimized token use in ConardLi/garden-skills.

- Repository: [ConardLi/garden-skills](https://github.com/ConardLi/garden-skills)
- Tags: how-to-guide
- Published: 2026-09-02

---

**The kb-retriever skill enables AI agents to answer questions over large, mixed-format knowledge bases by navigating hierarchical indexes, reading reference documentation before processing binary files, and performing progressive grep-based retrieval with windowed reads to minimize token consumption.**

The **kb-retriever** skill, available in the `ConardLi/garden-skills` repository, provides a token-efficient architecture for querying local directories containing Markdown, PDF, Excel, and text files. Unlike naive approaches that load entire documents into context, this skill uses a guided, iterative discovery process that preserves source citations while keeping computational costs low.

## Architecture of the kb-retriever Skill

The skill consists of four tightly-coupled components that work together to retrieve information without overwhelming the agent's context window.

### Hierarchical Index Navigation

The skill begins at the knowledge-base root (default `knowledge/`) and traverses a tree of [`data_structure.md`](https://github.com/ConardLi/garden-skills/blob/main/data_structure.md) files. Each index describes the purpose of its sub-directories and lists contained files, allowing the agent to prune irrelevant branches before examining content. According to [`skills/kb-retriever/SKILL.md`](https://github.com/ConardLi/garden-skills/blob/main/skills/kb-retriever/SKILL.md), this navigation ensures the agent focuses only on promising document paths.

### Learn-Before-Process Guardrails

When encountering PDF or Excel files, the skill **must first read reference documentation** located in [`references/pdf_reading.md`](https://github.com/ConardLi/garden-skills/blob/main/references/pdf_reading.md), [`references/excel_reading.md`](https://github.com/ConardLi/garden-skills/blob/main/references/excel_reading.md), or [`references/excel_analysis.md`](https://github.com/ConardLi/garden-skills/blob/main/references/excel_analysis.md). This mandatory step prevents blind processing of large binary files and guarantees optimal extraction tool selection, as documented in [`skills/kb-retriever/README.md`](https://github.com/ConardLi/garden-skills/blob/main/skills/kb-retriever/README.md).

### Progressive Retrieval Pipeline

Retrieval proceeds in three controlled steps:

1. **grep** – Locate keyword matches using precise `include` patterns to limit scope.
2. **Windowed reads** – Fetch only surrounding lines (`offset` + `limit`) of each match, avoiding full-file reads.
3. **Iterative refinement** – Repeat up to five rounds, tightening keyword sets until sufficient evidence is gathered.

This pipeline ensures minimal token usage even against large corpora.

### Per-Format Tool Strategy

The skill adopts distinct strategies per file type:

- **Markdown/text**: Direct `grep` followed by windowed `read_file` operations.
- **PDF**: Text extraction via `pdftotext` or `pdfplumber`, then grep the resulting `.txt`. For scanned PDFs lacking extractable text, use [`scripts/convert_pdf_to_images.py`](https://github.com/ConardLi/garden-skills/blob/main/scripts/convert_pdf_to_images.py).
- **Excel**: Preview schema with `pandas` using `nrows`, then filter rows/columns before grepping.

## Setting Up Your Knowledge Base

Create a hierarchical structure with index files to enable intelligent navigation:

```text
my-project/
├── knowledge/
│   ├── data_structure.md          # root index

│   ├── sales/
│   │   ├── data_structure.md
│   │   └── Q1_2023_sales.xlsx
│   └── manuals/
│       ├── data_structure.md
│       └── product_guide.pdf

```

Each [`data_structure.md`](https://github.com/ConardLi/garden-skills/blob/main/data_structure.md) should describe the directory's purpose and list contained files. The agent uses these indexes to build candidate file lists before attempting any content extraction.

## Query Execution Workflow

When processing a user question, the agent follows this sequence:

1. Determine the KB root and walk the [`data_structure.md`](https://github.com/ConardLi/garden-skills/blob/main/data_structure.md) hierarchy.
2. Build a candidate file list based on index relevance.
3. For PDF/Excel candidates, read the appropriate reference documentation first.
4. Extract or convert the file using the learned tool options.
5. Execute grep → windowed read → iterate (maximum 5 rounds).
6. Assemble the final answer with explicit citations (file name, line/page range).

As implemented in `ConardLi/garden-skills`, this flow ensures every step is gated and incremental.

## Handling PDF and Excel Files

For PDF queries, the agent consults [`references/pdf_reading.md`](https://github.com/ConardLi/garden-skills/blob/main/references/pdf_reading.md) before extraction:

```bash

# Extract text from pages 1-10 only

pdftotext -f 1 -l 10 manuals/product_guide.pdf /tmp/guide.txt

# Grep for the target term

grep -n "installation" /tmp/guide.txt | head

# Read a specific window (lines 100-150)

sed -n '100,150p' /tmp/guide.txt

```

For scanned PDFs without text layers, use the conversion script:

```python
python scripts/convert_pdf_to_images.py \
    --input manuals/scanned_manual.pdf \
    --output_dir /tmp/scanned_pages \
    --dpi 300

```

For Excel files, the agent reads [`references/excel_reading.md`](https://github.com/ConardLi/garden-skills/blob/main/references/excel_reading.md) and [`references/excel_analysis.md`](https://github.com/ConardLi/garden-skills/blob/main/references/excel_analysis.md), then uses pandas with row limits:

```python
import pandas as pd

# Preview schema only

df = pd.read_excel('sales/Q1_2023_sales.xlsx', nrows=10)
total = df['Revenue'].sum()

```

## Summary

- The **kb-retriever** skill uses hierarchical [`data_structure.md`](https://github.com/ConardLi/garden-skills/blob/main/data_structure.md) indexes to navigate large knowledge bases efficiently.
- It enforces **learn-before-process** guardrails, requiring agents to read reference docs in `references/` before handling PDFs or Excel files.
- **Progressive retrieval** combines grep, windowed reads, and up to five refinement rounds to minimize token consumption.
- Source citations (file paths and line ranges) are automatically included in final answers according to [`skills/kb-retriever/SKILL.md`](https://github.com/ConardLi/garden-skills/blob/main/skills/kb-retriever/SKILL.md).

## Frequently Asked Questions

### What file formats does the kb-retriever skill support?

The skill supports Markdown, plain text, PDF, and Excel files. For PDFs, it uses `pdftotext` or `pdfplumber` for text extraction, with a fallback to image conversion via [`scripts/convert_pdf_to_images.py`](https://github.com/ConardLi/garden-skills/blob/main/scripts/convert_pdf_to_images.py) for scanned documents. Excel files are processed using pandas with preview options to limit memory usage.

### How does the skill avoid loading entire files into context?

The skill implements **progressive retrieval**: it first uses grep to locate keyword matches, then performs **windowed reads** to fetch only specific line ranges (controlled by `offset` and `limit` parameters) around those matches. This process iterates up to five times, tightening search terms until sufficient evidence is found without ever loading complete files.

### Why must the agent read reference documents before processing PDFs or Excel files?

According to the architecture in [`skills/kb-retriever/README.md`](https://github.com/ConardLi/garden-skills/blob/main/skills/kb-retriever/README.md), reading [`references/pdf_reading.md`](https://github.com/ConardLi/garden-skills/blob/main/references/pdf_reading.md) or [`references/excel_reading.md`](https://github.com/ConardLi/garden-skills/blob/main/references/excel_reading.md) first prevents blind processing of large binary formats. This ensures the agent selects the correct extraction tool and applies optimal parameters like `nrows` for Excel previews, guaranteeing efficient and accurate data retrieval.

### Where is the skill configuration stored?

Skill metadata including name, description, and entry points resides in [`manifest.json`](https://github.com/ConardLi/garden-skills/blob/main/manifest.json) at the skill root. Detailed execution workflows are documented in [`SKILL.md`](https://github.com/ConardLi/garden-skills/blob/main/SKILL.md), while [`README.md`](https://github.com/ConardLi/garden-skills/blob/main/README.md) provides human-readable guides and best practices for implementation.