# Understanding the Retrieval Layers and Rules in kb-retriever

> Explore the four retrieval layers and hard-coded safety rules in kb-retriever for efficient, deterministic knowledge extraction from local repositories. Prevent full-file loading and optimize token usage.

- Repository: [ConardLi/garden-skills](https://github.com/ConardLi/garden-skills)
- Tags: deep-dive
- Published: 2026-08-31

---

**The kb-retriever skill implements a four-layer retrieval pipeline with hard-coded safety rules that prevent full-file loading and ensure token-efficient, deterministic knowledge extraction from local multi-format repositories.**

The **kb-retriever** skill in the [ConardLi/garden-skills](https://github.com/ConardLi/garden-skills) repository provides a disciplined architecture for querying local knowledge bases without exhausting LLM context windows. Understanding the **retrieval layers and rules in kb-retriever** is essential for developers building safe, scalable document processing agents. This skill operates through a multi-stage pipeline that strictly enforces format-specific tooling, hierarchical navigation, and bounded iteration loops.

## The Four Retrieval Layers in kb-retriever

The architecture is built around four distinct retrieval layers, each designed to minimize token consumption while maximizing answer accuracy.

### 1. Hierarchical Index Navigation

Before opening any document, the agent navigates a tree of [`data_structure.md`](https://github.com/ConardLi/garden-skills/blob/main/data_structure.md) files to identify the most relevant sub-directories for the query. This prevents blind scans of the entire corpus.

In the source implementation (lines 94-99), the skill enforces a strict rule: **only read directories that are explicitly indexed**. This hierarchical pre-filtering ensures the agent never wastes tokens examining irrelevant branches of the knowledge base.

### 2. Learn-Before-Process

When targeting PDF or Excel files, the skill first loads the appropriate reference guide from the `references/` directory—such as [`pdf_reading.md`](https://github.com/ConardLi/garden-skills/blob/main/pdf_reading.md) or [`excel_reading.md`](https://github.com/ConardLi/garden-skills/blob/main/excel_reading.md)—to learn the correct extraction tools and flags.

As coded in lines 100-107, the hard-coded rule states: **never process a PDF/Excel without first reading its reference guide**. This prevents tool misconfiguration and ensures the agent understands format-specific extraction parameters before touching binary files.

### 3. Progressive Retrieval

Instead of loading entire documents into context, the skill employs a **grep-first** strategy followed by windowed reads. The implementation (lines 117-123) uses `offset` and `limit` parameters to read only small windows around each match.

The governing rule is absolute: **never load whole files; always use windowed reads**. This layer is critical for maintaining token efficiency when processing large technical manuals or extensive Excel workbooks.

### 4. Per-Format Tool Strategy

Each file type triggers a specific extraction pipeline (lines 124-130):

- **Markdown/Text**: `grep` + windowed `read_file`
- **PDF**: `pdftotext` (or `pdfplumber`) → `grep`
- **Excel**: `pandas` with `nrows` → filtered reads

The rule mandates: **apply the correct tool per format; never use a generic approach**. This eliminates tool selection ambiguity and ensures deterministic extraction behavior across different content types.

## Iteration Control and Termination Rules

The retrieval process runs through a bounded iteration loop defined in lines 26-27 and 42-43. The system enforces a maximum of **5 rounds** of refinement, where each round re-applies the four layers above with updated keywords and candidate selection.

Early termination triggers include:

- Sufficient evidence gathered to answer the query (✓)
- Maximum round count reached (⏱️) — see line 42

This bounded approach prevents infinite loops and ensures predictable latency, even when queries are ambiguous or the knowledge base lacks relevant content.

## Key Implementation Files

Understanding the **retrieval layers and rules in kb-retriever** requires familiarity with these specific source files:

- [`skills/kb-retriever/SKILL.md`](https://github.com/ConardLi/garden-skills/blob/main/skills/kb-retriever/SKILL.md) — Front-matter definition declaring `name: kb-retriever`
- [`skills/kb-retriever/README.md`](https://github.com/ConardLi/garden-skills/blob/main/skills/kb-retriever/README.md) — Comprehensive documentation of the retrieval architecture
- [`skills/kb-retriever/references/pdf_reading.md`](https://github.com/ConardLi/garden-skills/blob/main/skills/kb-retriever/references/pdf_reading.md) — Extraction guidance for PDF processing
- [`skills/kb-retriever/references/excel_reading.md`](https://github.com/ConardLi/garden-skills/blob/main/skills/kb-retriever/references/excel_reading.md) — Pandas configuration for spreadsheet handling
- [`skills/kb-retriever/scripts/convert_pdf_to_images.py`](https://github.com/ConardLi/garden-skills/blob/main/skills/kb-retriever/scripts/convert_pdf_to_images.py) — Fallback for scanned PDFs when text extraction fails

## Usage Examples

Add the skill to your agent environment:

```bash

# Add the skill (once)

npx skills add ConardLi/garden-skills -s kb-retriever -a claude-code

```

Invoke the skill with a specific knowledge base path:

```json
{
  "skill": "kb-retriever",
  "question": "What are the main components of the Gantt chart in the project plan?",
  "kbPath": "./knowledge"
}

```

Direct JavaScript invocation:

```javascript
await agent.runSkill({
  name: "kb-retriever",
  args: {
    question: "Explain the design guidelines for the new branding assets."
    // kbPath defaults to ./knowledge if omitted
  }
});

```

Example agent prompt following the retrieval rules:

```markdown
You have the kb-retriever skill enabled.  
Answer the following using the skill:

> Summarize the key milestones from the file `project-plan.pdf` in the
> `project-management/` directory.

Remember: First read `references/pdf_reading.md` to learn the extraction
tool, then use `grep` to locate the term "milestone", finally read the
matching window.

```

## Summary

- **Four retrieval layers** govern the kb-retriever skill: Hierarchical Index Navigation, Learn-Before-Process, Progressive Retrieval, and Per-Format Tool Strategy
- **Hard-coded rules** prevent full-file loading and mandate format-specific tooling
- **Bounded iteration** limits retrieval to 5 rounds with early termination on success
- **Reference-gated processing** ensures agents learn extraction methods before touching PDFs or Excel files
- **Windowed reads** replace whole-file loading to maintain token efficiency

## Frequently Asked Questions

### What is the maximum number of retrieval rounds in kb-retriever?

The skill enforces a hard limit of **5 rounds** (lines 26-27, 42-43). Each round refines keywords and reapplies the four retrieval layers until sufficient evidence is found or the maximum count is reached.

### Why does kb-retriever require reading reference files before processing PDFs?

According to lines 100-107 of the implementation, the **Learn-Before-Process** rule prevents tool misconfiguration. By first reading [`references/pdf_reading.md`](https://github.com/ConardLi/garden-skills/blob/main/references/pdf_reading.md) or [`references/excel_reading.md`](https://github.com/ConardLi/garden-skills/blob/main/references/excel_reading.md), the agent learns the correct extraction flags and tools, ensuring deterministic processing of binary formats.

### How does kb-retriever handle large files without exceeding token limits?

The **Progressive Retrieval** layer (lines 117-123) uses a **grep-first** approach followed by windowed `read_file` calls using `offset` and `limit` parameters. The rule **never load whole files** ensures only relevant text windows enter the LLM context.

### Where are the retrieval rules documented in the source code?

The primary documentation resides in [`skills/kb-retriever/README.md`](https://github.com/ConardLi/garden-skills/blob/main/skills/kb-retriever/README.md), while the enforcement logic is implemented in the main skill script. Layer-specific behaviors are indexed at lines 94-130, and iteration controls appear at lines 26-27 and 42-43.