What Types of Local Files Can the kb-Retriever Skill Process?

The kb-retriever skill processes three local file types—Markdown or plain-text files, PDF documents, and Excel spreadsheets—using a token-efficient architecture that reads reference guides before processing and retrieves only relevant windowed segments rather than full files.

The kb-retriever skill in the ConardLi/garden-skills repository implements a disciplined approach to local file processing. Instead of loading entire documents into context, it employs hierarchical navigation and progressive retrieval to minimize token consumption while maintaining accuracy across supported formats.

Core File Types Supported

The skill handles three primary categories of local files through distinct processing pipelines defined in the /skills/kb-retriever/ directory.

Markdown and Plain-Text Files

For .md, .txt, and similar text formats, the skill uses grep to locate relevant keywords across the knowledge base. It then invokes read_file with offset and limit parameters to extract only the matching window, preventing whole-file loads. This approach ensures constant-time retrieval regardless of file size.

PDF Documents

When processing .pdf files, the skill first consults /skills/kb-retriever/references/pdf_reading.md to identify the appropriate extraction toolchain—typically pdftotext or pdfplumber. The PDF converts to a temporary text file, undergoes keyword-based grep filtering, and returns only the relevant text segments. For scanned PDFs where text extraction fails, the system falls back to /skills/kb-retriever/scripts/convert_pdf_to_images.py.

Excel Spreadsheets

For .xlsx and .xls files, the skill reads /skills/kb-retriever/references/excel_reading.md and /skills/kb-retriever/references/excel_analysis.md to learn column schemas and optimal pandas parameters like nrows and dtype. It performs filtered reads using these configurations, retrieving specific rows rather than loading entire spreadsheets into memory.

Processing Architecture

The kb-retriever skill implements three architectural constraints to ensure safe, efficient file handling.

Hierarchical Index Navigation

Before touching any target file, the skill walks data_structure.md files to prune the candidate set. This navigation reduces the search space and prevents unnecessary file system operations.

Learn-Before-Process Protocol

For non-text formats, the skill must read the associated reference markdown before processing the target file. This requirement, documented in section 2 of /skills/kb-retriever/README.md, ensures the skill selects appropriate tooling and understands format-specific safety constraints before execution.

Progressive Retrieval

The skill never loads complete files into context. Instead, it retrieves small, relevant windows determined by the initial grep search or pandas filtering criteria. This progressive approach caps token consumption while preserving answer quality.

Extending Support for Additional Formats

The architecture supports extensibility through the references/*.md file pattern. To add a new format, create a corresponding reference document describing the required toolchain and safe extraction methods. Until a reference document exists for a file type, the skill refuses processing to prevent uncontrolled token consumption.

Practical Usage Examples

The following examples demonstrate how to invoke the kb-retriever skill for each supported file type.

Querying Markdown Knowledge Bases

{
  "skill": "kb-retriever",
  "input": "What are the main steps for onboarding new team members?",
  "kb_path": "./knowledge"
}

This invocation searches markdown files under knowledge/, reads matching snippets, and returns answers with source citations.

Processing PDF Documents

{
  "skill": "kb-retriever",
  "input": "Summarize the privacy policy in the PDF file `privacy.pdf`",
  "kb_path": "./documents"
}

The skill reads references/pdf_reading.md to select the correct pdftotext command, extracts the PDF to plain text, and performs keyword-based retrieval.

Analyzing Excel Data

{
  "skill": "kb-retriever",
  "input": "What were the sales figures for Q2 2023?",
  "kb_path": "./reports"
}

The skill consults references/excel_reading.md to learn column structures, then uses pandas with filtered reads to retrieve only relevant rows.

Summary

  • The kb-retriever skill supports three file types: Markdown/plain-text, PDF, and Excel.
  • Processing relies on reference-guided extraction: the skill reads format-specific guides before handling non-text files.
  • Progressive retrieval via grep and windowed read_file calls ensures token-efficient processing without loading complete documents.
  • New formats require a references/*.md file; unsupported formats are rejected to prevent token overflow.
  • Source configuration resides in skills/kb-retriever/SKILL.md with implementation details in skills/kb-retriever/README.md.

Frequently Asked Questions

Can the kb-retriever skill process Word documents or CSV files?

No. According to the source code in skills/kb-retriever/README.md, the skill only processes Markdown, PDF, and Excel files. To handle Word documents or CSVs, you must first create a corresponding reference file in skills/kb-retriever/references/ that defines the extraction toolchain and safety parameters.

How does the skill prevent token overflow when processing large PDFs?

The skill implements a learn-before-process protocol using references/pdf_reading.md to select efficient tools like pdftotext or pdfplumber. It converts the PDF to text, runs grep to identify relevant sections, and uses read_file with offset and limit parameters to retrieve only small windows of text, never loading the full document.

What happens if a PDF contains scanned images instead of text?

When text extraction fails, the skill falls back to scripts/convert_pdf_to_images.py to convert scanned pages to images for alternative processing. However, the primary workflow favors text-based extraction via pdftotext or pdfplumber as defined in the PDF reading reference guide.

Is the kb-retriever skill limited to local files only?

Yes. The architecture as defined in skills/kb-retriever/SKILL.md targets local file systems via the kb_path parameter. The skill walks local directories, reads local reference files, and processes local documents using system tools like grep and file-specific converters.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →