How Agent Skills Reduce Tool Selection Burden in LLM Systems
Agent Skills use a progressive disclosure mechanism that transforms tool selection from a combinatorial search problem into a lightweight knowledge retrieval task, keeping system prompts compact while loading full capabilities on demand.
The bojieli/ai-agent-book repository introduces Agent Skills as an architectural pattern designed to solve the tool selection burden that paralyzes large language models when presented with extensive tool catalogs. By implementing a three-layer progressive disclosure system, this approach collapses the complex "which tool" decision into a simple string-matching operation, dramatically reducing token consumption and reasoning overhead.
The Problem: Context Explosion in Multi-Tool Agents
When an agent has access to hundreds of specialized tools, the system prompt must include every tool's name, description, and parameter schema. This creates several systemic issues:
- KV-cache misses: Long prompts exhaust cache efficiency and increase latency
- Reasoning degradation: The LLM must engage in complex combinatorial reasoning to select the correct tool from a massive catalog
- Token budget exhaustion: Tool descriptions consume the context window, leaving less room for actual task execution
According to Chapter 2 of the book, as prompts grow, "the static system prompt becomes infeasible" (lines 736-751 in book/chapter2.md).
How Agent Skills Reduce Tool Selection Burden Through Progressive Disclosure
The implementation in chapter2/agent-skills-ppt/demo.py demonstrates how Agent Skills eliminate selection complexity through four distinct layers of abstraction:
Layer 1: Thin Catalog Initialization
When an Agent instantiates, it receives only a compact directory containing each Skill's name and a short description (approximately a few hundred tokens). The function scan_skill_catalog() (lines 9-13) builds this lightweight index without loading concrete procedures. This keeps the system prompt short and preserves KV-Cache hits.
Layer 2: On-Demand Skill Materialization
When a task requires specific capabilities, the Agent invokes the read_skill tool. The full SKILL.md—containing the workflow, prompts, and parameters—is then streamed into the context (lines 53-58). This represents the second layer of knowledge, loaded only when explicitly needed.
Layer 3: Granular Asset Retrieval
For auxiliary assets such as reference documents or bundled scripts, the Agent uses read_skill_file to fetch only the specific file required (lines 17-23). This third layer prevents context pollution by excluding irrelevant skill resources from the active context window.
Layer 4: Isolated Execution Environment
The run_skill_script tool (lines 33-41) executes the script belonging to that specific Skill, guaranteeing the Agent never attempts to invoke an unknown tool or stray code. This isolation ensures that the selection problem remains confined to the skill level rather than individual tool parameters.
Implementation Example: The Three-Layer Workflow
The following Python code demonstrates the progressive disclosure mechanism used in the repository:
from pathlib import Path
import json
# Layer 1: Scan the thin catalog (first layer)
def scan_skill_catalog() -> dict:
catalog = {}
for md in sorted(Path("skills").glob("*/SKILL.md")):
meta = {}
if md.read_text().startswith("---"):
end = md.read_text().find("---", 3)
for line in md.read_text()[:end].splitlines():
if ":" in line:
k, v = line.split(":", 1)
meta[k.strip()] = v.strip()
name = meta.get("name") or md.parent.name
catalog[name] = {"description": meta.get("description", ""), "dir": md.parent}
return catalog
catalog = scan_skill_catalog()
print("Thin catalog →", list(catalog)[:3]) # only names & short desc.
# Layer 2: Load the full Skill when needed (second layer)
def read_skill(name: str) -> str:
info = catalog[name]
return (info["dir"] / "SKILL.md").read_text(encoding="utf-8")
full_skill_md = read_skill("pptx")
print("\n--- Loaded SKILL.md ---\n", full_skill_md[:200])
# Layer 3: Retrieve a helper file (third layer)
def read_skill_file(name: str, rel_path: str) -> str:
info = catalog[name]
target = (info["dir"] / rel_path).resolve()
assert str(target).startswith(str(info["dir"].resolve()))
return target.read_text(encoding="utf-8")
ref = read_skill_file("pptx", "reference.md")
print("\n--- reference.md snippet ---\n", ref[:150])
# Layer 4: Execute the bundled script
def run_skill_script(name: str, script: str, payload: dict, out_path: Path) -> dict:
info = catalog[name]
scripts_dir = (info["dir"] / "scripts").resolve()
script_path = (scripts_dir / script).resolve()
assert script_path.is_relative_to(scripts_dir)
import importlib.util
spec = importlib.util.spec_from_file_location("bundled", script_path)
mod = importlib.util.module_from_spec(spec)
spec.loader.exec_module(mod)
return mod.build_presentation(payload, str(out_path))
result = run_skill_script(
"pptx",
"generate_pptx.py",
{"title": "Demo", "slides": [{"title": "Slide 1"}]},
Path("output/demo.pptx")
)
Architectural Impact: From Tool Selection to Knowledge Retrieval
Chapter 4 of the book (lines 142-150 in book/chapter4.md) contrasts MCP (Model Context Protocol), which focuses on interoperability, with Skills, which address choice overload.
Agent Skills reduce tool selection burden by replacing a large toolbox of specialized tools with a small set of generic executors plus on-demand knowledge. This architectural shift transforms the problem from "which of these 500 tools should I use?" into "does any Skill match the current intent?"—a simple retrieval operation requiring only a few tokens.
The result is a dramatic reduction in the token budget spent on tool descriptions, elimination of combinatorial explosion in tool-selection reasoning, and preserved LLM capacity for actual task execution.
Summary
- Progressive disclosure collapses tool selection into lightweight knowledge retrieval by exposing only Skill names initially, then loading full implementations on demand.
- The four-layer architecture in
demo.py(thin catalog →read_skill→read_skill_file→run_skill_script) isolates complexity and preserves KV-cache efficiency. - Context window preservation is achieved by keeping concrete tool definitions out of the system prompt until explicitly needed.
- Execution isolation via
run_skill_scriptensures the Agent never invokes unknown tools, reducing the selection surface area to the Skill level only. - Compared to MCP, Skills prioritize choice reduction over interoperability, solving the scaling problem of large tool catalogs.
Frequently Asked Questions
How do Agent Skills differ from MCP (Model Context Protocol)?
According to book/chapter4.md (lines 142-150), MCP focuses on interoperability between systems, while Agent Skills specifically target choice overload. Skills replace extensive tool listings with compact, on-demand knowledge retrieval, whereas MCP standardizes how tools are exposed without reducing the selection burden.
What specific problem does progressive disclosure solve in LLM agents?
Progressive disclosure solves the context explosion that occurs when system prompts must include hundreds of tool schemas and descriptions. By loading only the Skill name initially (as implemented in scan_skill_catalog), the mechanism prevents KV-cache misses and preserves reasoning capacity for the actual task rather than tool selection.
How does the thin catalog mechanism preserve KV-cache efficiency?
The thin catalog contains only Skill names and brief descriptions (approximately a few hundred tokens) rather than full tool implementations. As noted in chapter2/agent-skills-ppt/demo.py (lines 9-13), this compact representation stays resident in the KV-cache, while the full SKILL.md content is loaded only when read_skill is invoked.
Can Agent Skills handle multiple skills in a single conversation?
Yes. The Agent can invoke read_skill multiple times during a single session to materialize different capabilities as needed. The run_skill_script function ensures each Skill's execution remains isolated, preventing cross-contamination while allowing the Agent to compose workflows across multiple Skills through sequential retrieval.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →