# What Types of Local Files Can the kb-Retriever Skill Process?

> Discover the local file types the kb-retriever skill processes, including Markdown, PDF, and Excel. Learn how its efficient architecture retrieves relevant data.

- Repository: [ConardLi/garden-skills](https://github.com/ConardLi/garden-skills)
- Tags: how-to-guide
- Published: 2026-08-28

---

**The kb-retriever skill processes three local file types—Markdown or plain-text files, PDF documents, and Excel spreadsheets—using a token-efficient architecture that reads reference guides before processing and retrieves only relevant windowed segments rather than full files.**

The kb-retriever skill in the ConardLi/garden-skills repository implements a disciplined approach to local file processing. Instead of loading entire documents into context, it employs hierarchical navigation and progressive retrieval to minimize token consumption while maintaining accuracy across supported formats.

## Core File Types Supported

The skill handles three primary categories of local files through distinct processing pipelines defined in the `/skills/kb-retriever/` directory.

### Markdown and Plain-Text Files

For `.md`, `.txt`, and similar text formats, the skill uses `grep` to locate relevant keywords across the knowledge base. It then invokes `read_file` with **offset** and **limit** parameters to extract only the matching window, preventing whole-file loads. This approach ensures constant-time retrieval regardless of file size.

### PDF Documents

When processing `.pdf` files, the skill first consults [`/skills/kb-retriever/references/pdf_reading.md`](https://github.com/ConardLi/garden-skills/blob/main//skills/kb-retriever/references/pdf_reading.md) to identify the appropriate extraction toolchain—typically `pdftotext` or `pdfplumber`. The PDF converts to a temporary text file, undergoes keyword-based `grep` filtering, and returns only the relevant text segments. For scanned PDFs where text extraction fails, the system falls back to [`/skills/kb-retriever/scripts/convert_pdf_to_images.py`](https://github.com/ConardLi/garden-skills/blob/main//skills/kb-retriever/scripts/convert_pdf_to_images.py).

### Excel Spreadsheets

For `.xlsx` and `.xls` files, the skill reads [`/skills/kb-retriever/references/excel_reading.md`](https://github.com/ConardLi/garden-skills/blob/main//skills/kb-retriever/references/excel_reading.md) and [`/skills/kb-retriever/references/excel_analysis.md`](https://github.com/ConardLi/garden-skills/blob/main//skills/kb-retriever/references/excel_analysis.md) to learn column schemas and optimal `pandas` parameters like `nrows` and `dtype`. It performs filtered reads using these configurations, retrieving specific rows rather than loading entire spreadsheets into memory.

## Processing Architecture

The kb-retriever skill implements three architectural constraints to ensure safe, efficient file handling.

### Hierarchical Index Navigation

Before touching any target file, the skill walks [`data_structure.md`](https://github.com/ConardLi/garden-skills/blob/main/data_structure.md) files to prune the candidate set. This navigation reduces the search space and prevents unnecessary file system operations.

### Learn-Before-Process Protocol

For non-text formats, the skill must read the associated reference markdown before processing the target file. This requirement, documented in section 2 of [`/skills/kb-retriever/README.md`](https://github.com/ConardLi/garden-skills/blob/main//skills/kb-retriever/README.md), ensures the skill selects appropriate tooling and understands format-specific safety constraints before execution.

### Progressive Retrieval

The skill never loads complete files into context. Instead, it retrieves small, relevant windows determined by the initial `grep` search or pandas filtering criteria. This progressive approach caps token consumption while preserving answer quality.

## Extending Support for Additional Formats

The architecture supports extensibility through the `references/*.md` file pattern. To add a new format, create a corresponding reference document describing the required toolchain and safe extraction methods. Until a reference document exists for a file type, the skill refuses processing to prevent uncontrolled token consumption.

## Practical Usage Examples

The following examples demonstrate how to invoke the kb-retriever skill for each supported file type.

### Querying Markdown Knowledge Bases

```json
{
  "skill": "kb-retriever",
  "input": "What are the main steps for onboarding new team members?",
  "kb_path": "./knowledge"
}

```

This invocation searches markdown files under `knowledge/`, reads matching snippets, and returns answers with source citations.

### Processing PDF Documents

```json
{
  "skill": "kb-retriever",
  "input": "Summarize the privacy policy in the PDF file `privacy.pdf`",
  "kb_path": "./documents"
}

```

The skill reads [`references/pdf_reading.md`](https://github.com/ConardLi/garden-skills/blob/main/references/pdf_reading.md) to select the correct `pdftotext` command, extracts the PDF to plain text, and performs keyword-based retrieval.

### Analyzing Excel Data

```json
{
  "skill": "kb-retriever",
  "input": "What were the sales figures for Q2 2023?",
  "kb_path": "./reports"
}

```

The skill consults [`references/excel_reading.md`](https://github.com/ConardLi/garden-skills/blob/main/references/excel_reading.md) to learn column structures, then uses `pandas` with filtered reads to retrieve only relevant rows.

## Summary

- The kb-retriever skill supports three file types: **Markdown/plain-text**, **PDF**, and **Excel**.
- Processing relies on **reference-guided extraction**: the skill reads format-specific guides before handling non-text files.
- **Progressive retrieval** via `grep` and windowed `read_file` calls ensures token-efficient processing without loading complete documents.
- New formats require a `references/*.md` file; unsupported formats are rejected to prevent token overflow.
- Source configuration resides in [`skills/kb-retriever/SKILL.md`](https://github.com/ConardLi/garden-skills/blob/main/skills/kb-retriever/SKILL.md) with implementation details in [`skills/kb-retriever/README.md`](https://github.com/ConardLi/garden-skills/blob/main/skills/kb-retriever/README.md).

## Frequently Asked Questions

### Can the kb-retriever skill process Word documents or CSV files?

No. According to the source code in [`skills/kb-retriever/README.md`](https://github.com/ConardLi/garden-skills/blob/main/skills/kb-retriever/README.md), the skill only processes Markdown, PDF, and Excel files. To handle Word documents or CSVs, you must first create a corresponding reference file in `skills/kb-retriever/references/` that defines the extraction toolchain and safety parameters.

### How does the skill prevent token overflow when processing large PDFs?

The skill implements a learn-before-process protocol using [`references/pdf_reading.md`](https://github.com/ConardLi/garden-skills/blob/main/references/pdf_reading.md) to select efficient tools like `pdftotext` or `pdfplumber`. It converts the PDF to text, runs `grep` to identify relevant sections, and uses `read_file` with offset and limit parameters to retrieve only small windows of text, never loading the full document.

### What happens if a PDF contains scanned images instead of text?

When text extraction fails, the skill falls back to [`scripts/convert_pdf_to_images.py`](https://github.com/ConardLi/garden-skills/blob/main/scripts/convert_pdf_to_images.py) to convert scanned pages to images for alternative processing. However, the primary workflow favors text-based extraction via `pdftotext` or `pdfplumber` as defined in the PDF reading reference guide.

### Is the kb-retriever skill limited to local files only?

Yes. The architecture as defined in [`skills/kb-retriever/SKILL.md`](https://github.com/ConardLi/garden-skills/blob/main/skills/kb-retriever/SKILL.md) targets local file systems via the `kb_path` parameter. The skill walks local directories, reads local reference files, and processes local documents using system tools like `grep` and file-specific converters.