# How Agent Skills Reduce Tool Selection Burden in LLM Systems

> Discover how agent skills simplify LLM tool selection from complex search to efficient knowledge retrieval. Keep prompts compact and capabilities ready on demand.

- Repository: [Bojie Li/ai-agent-book](https://github.com/bojieli/ai-agent-book)
- Tags: deep-dive
- Published: 2026-08-24

---

**Agent Skills use a progressive disclosure mechanism that transforms tool selection from a combinatorial search problem into a lightweight knowledge retrieval task, keeping system prompts compact while loading full capabilities on demand.**

The `bojieli/ai-agent-book` repository introduces **Agent Skills** as an architectural pattern designed to solve the tool selection burden that paralyzes large language models when presented with extensive tool catalogs. By implementing a three-layer progressive disclosure system, this approach collapses the complex "which tool" decision into a simple string-matching operation, dramatically reducing token consumption and reasoning overhead.

## The Problem: Context Explosion in Multi-Tool Agents

When an agent has access to hundreds of specialized tools, the system prompt must include every tool's name, description, and parameter schema. This creates several systemic issues:

- **KV-cache misses**: Long prompts exhaust cache efficiency and increase latency
- **Reasoning degradation**: The LLM must engage in complex combinatorial reasoning to select the correct tool from a massive catalog
- **Token budget exhaustion**: Tool descriptions consume the context window, leaving less room for actual task execution

According to Chapter 2 of the book, as prompts grow, "the static system prompt becomes infeasible" (lines 736-751 in [`book/chapter2.md`](https://github.com/bojieli/ai-agent-book/blob/main/book/chapter2.md)).

## How Agent Skills Reduce Tool Selection Burden Through Progressive Disclosure

The implementation in [`chapter2/agent-skills-ppt/demo.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter2/agent-skills-ppt/demo.py) demonstrates how Agent Skills eliminate selection complexity through four distinct layers of abstraction:

### Layer 1: Thin Catalog Initialization

When an Agent instantiates, it receives only a compact directory containing each Skill's `name` and a short `description` (approximately a few hundred tokens). The function `scan_skill_catalog()` (lines 9-13) builds this lightweight index without loading concrete procedures. This keeps the system prompt short and preserves KV-Cache hits.

### Layer 2: On-Demand Skill Materialization

When a task requires specific capabilities, the Agent invokes the `read_skill` tool. The full [`SKILL.md`](https://github.com/bojieli/ai-agent-book/blob/main/SKILL.md)—containing the workflow, prompts, and parameters—is then streamed into the context (lines 53-58). This represents the second layer of knowledge, loaded only when explicitly needed.

### Layer 3: Granular Asset Retrieval

For auxiliary assets such as reference documents or bundled scripts, the Agent uses `read_skill_file` to fetch only the specific file required (lines 17-23). This third layer prevents context pollution by excluding irrelevant skill resources from the active context window.

### Layer 4: Isolated Execution Environment

The `run_skill_script` tool (lines 33-41) executes the script belonging to that specific Skill, guaranteeing the Agent never attempts to invoke an unknown tool or stray code. This isolation ensures that the selection problem remains confined to the skill level rather than individual tool parameters.

## Implementation Example: The Three-Layer Workflow

The following Python code demonstrates the progressive disclosure mechanism used in the repository:

```python
from pathlib import Path
import json

# Layer 1: Scan the thin catalog (first layer)

def scan_skill_catalog() -> dict:
    catalog = {}
    for md in sorted(Path("skills").glob("*/SKILL.md")):
        meta = {}
        if md.read_text().startswith("---"):
            end = md.read_text().find("---", 3)
            for line in md.read_text()[:end].splitlines():
                if ":" in line:
                    k, v = line.split(":", 1)
                    meta[k.strip()] = v.strip()
        name = meta.get("name") or md.parent.name
        catalog[name] = {"description": meta.get("description", ""), "dir": md.parent}
    return catalog

catalog = scan_skill_catalog()
print("Thin catalog →", list(catalog)[:3])   # only names & short desc.

# Layer 2: Load the full Skill when needed (second layer)

def read_skill(name: str) -> str:
    info = catalog[name]
    return (info["dir"] / "SKILL.md").read_text(encoding="utf-8")

full_skill_md = read_skill("pptx")
print("\n--- Loaded SKILL.md ---\n", full_skill_md[:200])

# Layer 3: Retrieve a helper file (third layer)

def read_skill_file(name: str, rel_path: str) -> str:
    info = catalog[name]
    target = (info["dir"] / rel_path).resolve()
    assert str(target).startswith(str(info["dir"].resolve()))
    return target.read_text(encoding="utf-8")

ref = read_skill_file("pptx", "reference.md")
print("\n--- reference.md snippet ---\n", ref[:150])

# Layer 4: Execute the bundled script

def run_skill_script(name: str, script: str, payload: dict, out_path: Path) -> dict:
    info = catalog[name]
    scripts_dir = (info["dir"] / "scripts").resolve()
    script_path = (scripts_dir / script).resolve()
    assert script_path.is_relative_to(scripts_dir)
    import importlib.util
    spec = importlib.util.spec_from_file_location("bundled", script_path)
    mod = importlib.util.module_from_spec(spec)
    spec.loader.exec_module(mod)
    return mod.build_presentation(payload, str(out_path))

result = run_skill_script(
    "pptx",
    "generate_pptx.py",
    {"title": "Demo", "slides": [{"title": "Slide 1"}]},
    Path("output/demo.pptx")
)

```

## Architectural Impact: From Tool Selection to Knowledge Retrieval

Chapter 4 of the book (lines 142-150 in [`book/chapter4.md`](https://github.com/bojieli/ai-agent-book/blob/main/book/chapter4.md)) contrasts **MCP** (Model Context Protocol), which focuses on interoperability, with **Skills**, which address choice overload.

**Agent Skills reduce tool selection burden** by replacing a large toolbox of specialized tools with a small set of generic executors plus on-demand knowledge. This architectural shift transforms the problem from "which of these 500 tools should I use?" into "does any Skill match the current intent?"—a simple retrieval operation requiring only a few tokens.

The result is a dramatic reduction in the token budget spent on tool descriptions, elimination of combinatorial explosion in tool-selection reasoning, and preserved LLM capacity for actual task execution.

## Summary

- **Progressive disclosure** collapses tool selection into lightweight knowledge retrieval by exposing only Skill names initially, then loading full implementations on demand.
- The four-layer architecture in [`demo.py`](https://github.com/bojieli/ai-agent-book/blob/main/demo.py) (thin catalog → `read_skill` → `read_skill_file` → `run_skill_script`) isolates complexity and preserves KV-cache efficiency.
- **Context window preservation** is achieved by keeping concrete tool definitions out of the system prompt until explicitly needed.
- **Execution isolation** via `run_skill_script` ensures the Agent never invokes unknown tools, reducing the selection surface area to the Skill level only.
- Compared to MCP, Skills prioritize **choice reduction** over interoperability, solving the scaling problem of large tool catalogs.

## Frequently Asked Questions

### How do Agent Skills differ from MCP (Model Context Protocol)?

According to [`book/chapter4.md`](https://github.com/bojieli/ai-agent-book/blob/main/book/chapter4.md) (lines 142-150), MCP focuses on interoperability between systems, while Agent Skills specifically target choice overload. Skills replace extensive tool listings with compact, on-demand knowledge retrieval, whereas MCP standardizes how tools are exposed without reducing the selection burden.

### What specific problem does progressive disclosure solve in LLM agents?

Progressive disclosure solves the **context explosion** that occurs when system prompts must include hundreds of tool schemas and descriptions. By loading only the Skill name initially (as implemented in `scan_skill_catalog`), the mechanism prevents KV-cache misses and preserves reasoning capacity for the actual task rather than tool selection.

### How does the thin catalog mechanism preserve KV-cache efficiency?

The thin catalog contains only Skill names and brief descriptions (approximately a few hundred tokens) rather than full tool implementations. As noted in [`chapter2/agent-skills-ppt/demo.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter2/agent-skills-ppt/demo.py) (lines 9-13), this compact representation stays resident in the KV-cache, while the full [`SKILL.md`](https://github.com/bojieli/ai-agent-book/blob/main/SKILL.md) content is loaded only when `read_skill` is invoked.

### Can Agent Skills handle multiple skills in a single conversation?

Yes. The Agent can invoke `read_skill` multiple times during a single session to materialize different capabilities as needed. The `run_skill_script` function ensures each Skill's execution remains isolated, preventing cross-contamination while allowing the Agent to compose workflows across multiple Skills through sequential retrieval.