What Are Agent Skills? A Modular Approach to AI Agent Context Engineering
Agent Skills are a modular, on-demand knowledge-pack mechanism that lets an AI Agent load domain-specific capabilities only when needed, instead of overloading a static system prompt with every possible instruction.
In context engineering, bloated system prompts create two critical problems: token waste from irrelevant text and attention dilution that reduces model focus. Agent Skills solve this through progressive disclosure—the Agent sees only a lightweight catalog of available skills until task matching triggers full skill loading. This design pattern, detailed in Bojie Li's AI Agent Book, keeps static prefixes short, preserves KV-Cache hits, and enables domain scaling without repeated token costs.
Why Agent Skills Matter in Context Engineering
Traditional monolithic prompts fail as Agents gain capabilities. Every new domain—PowerPoint generation, database querying, Chinese prose rewriting—adds hundreds or thousands of tokens to the system prompt, most wasted on any given task.
The book/chapter2.md file identifies this as a fundamental bottleneck. Agent Skills replace this static approach with a two-layer architecture that defers heavy payload loading until explicitly required.
The Two-Layer Skill Architecture
Agent Skills organize capability into Layer 1 (Metadata) and Layer 2 (Core Flow).
Layer 1: Skill Catalog (Metadata)
At startup, the framework scans all SKILL.md files and extracts only their YAML front-matter. This creates a compact catalog—typically a few hundred tokens—that sits in the persistent context.
The skill_manager.py file (lines 75-87) implements this extraction:
# From chapter9/ai-style-skill/skill_manager.py
# Renders SKILL.md metadata without loading full skill text
def render_skill_md(skill_data: dict) -> str:
"""
Extracts YAML front-matter name/description for catalog display.
Full skill definition deferred until routing decision.
"""
return f"{skill_data['name']}: {skill_data['description']}"
A minimal SKILL.md demonstrates this structure:
---
name: my-summarizer
description: "When the user asks for a concise summary of a long document, load this skill."
---
# Summarizer Skill
## When to Load
When the user says "Summarize the following text" or provides a URL longer than 500 words.
## Rules
### Rule 1: Extract key points
- **Definition:** Identify the three most important sentences.
- **Detector:** LLM-based semantic ranking.
- **Bad example:** "The text is about cats."
- **Good example:** "Cats are small, domesticated mammals."
Layer 2: Full Skill Definition
The content after the YAML front-matter—rules, examples, tool scripts—remains on disk until the Agent's routing logic matches the skill description to the current task. Then and only then does the full payload inject into context.
Progressive Disclosure and KV-Cache Efficiency
The progressive disclosure mechanism delivers measurable performance benefits:
- Low upfront cost: Only the catalog (metadata) occupies the context window initially
- Cache preservation: The heavy payload appends after the cached static prefix, so existing KV-Cache entries remain valid
- Addition-at-end safety: As noted in
book/chapter2.md, appending content does not break caching; prepending or interleaving would
This matters at scale. An Agent with 50 skills averaging 2,000 tokens each would need 100,000+ tokens in a monolithic prompt. With Agent Skills, only the ~500-token catalog stays resident.
Skill Structure: Files, Scripts, and Safety
A Skill is more than text. The PPTX Skill example (chapter2/agent-skills-ppt/) shows a complete bundle:
skills/pptx/
├── SKILL.md # Definition and metadata
└── scripts/
├── generate_pptx.py # Python helper
└── html2pptx.js # JavaScript converter
The chapter2/agent-skills-ppt/demo.py driver orchestrates execution:
- Catalog scan → discovers
pptxskill metadata - Task matching →
/pptxcommand triggers skill loading - Full injection →
SKILL.mdcontent plus helper scripts enter context - Tool execution → LLM-driven workflow assembles the presentation
# Trigger PPTX generation via skill loading
python -m chapter2.agent-skills-ppt.demo /pptx \
--paper attention-is-all-you-need.pdf \
--output deck.pptx
Security consideration: Skills carry executable code. The book emphasizes that unknown skills require audit before loading—treat them as you would unreviewed scripts.
Implementing Skill Management in Python
The repository's skill_manager.py provides utilities for skill manipulation:
from skill_manager import write_skill, render_skill_md
# Programmatic skill creation
skill_path = write_skill([{
"id": "my-summarizer-001",
"name": "Summarizer",
"definition": "Extract three key sentences from the input.",
"detector": {"type": "llm"},
"bad_example": "The text talks about cats.",
"good_example": "Cats are small, domesticated mammals.",
"scope": ["document"],
"source_ids": ["user-feedback-123"]
}])
print(f"Skill written to {skill_path}")
Key Source Files for Deep Dives
| File | Purpose |
|---|---|
book/chapter2.md |
Core motivation and progressive-disclosure design |
chapter9/ai-style-skill/skill_manager.py |
Skill merging, pruning, and SKILL.md rendering |
chapter2/agent-skills-ppt/README.md |
End-to-end PPTX Skill walkthrough |
chapter2/agent-skills-ppt/skills/pptx/scripts/generate_pptx.py |
Bundled tooling example |
chapter2/agent-skills-ppt/demo.py |
Full orchestration driver |
Summary
- Agent Skills replace monolithic prompts with modular, load-on-demand capability bundles
- Two-layer architecture separates lightweight metadata (always present) from heavy payload (loaded per-task)
- Progressive disclosure minimizes token costs and preserves KV-Cache efficiency
- SKILL.md format uses YAML front-matter for routing and markdown body for full definition
- Security model treats skills as auditable external code, not implicit system extensions
Frequently Asked Questions
How do Agent Skills differ from function calling or tool use?
Function calling exposes discrete operations an Agent can invoke; Agent Skills are self-contained capability packages that include instructions, examples, and helper tooling. A Skill might internally use multiple function calls, but its defining feature is the bundled knowledge that only loads when contextually relevant.
What happens if multiple Skills match a single task?
The routing layer—implemented in skill_manager.py—uses the description field as a semantic routing cue. Conflicts resolve through explicit priority markers in the YAML metadata or LLM-based arbitration. The design encourages orthogonal Skill descriptions to minimize overlap.
Can Skills depend on or compose other Skills?
The base specification focuses on single-Skill loading, but chapter9/ai-style-skill/skill_manager.py supports Skill merging for compound operations. A meta-Skill can reference multiple base Skills in its definition, triggering sequential or parallel loading.
Why YAML front-matter instead of JSON or TOML?
YAML front-matter permits human-readable metadata that renders cleanly in markdown viewers while remaining machine-parseable. The --- delimiter convention integrates with static site generators and documentation workflows, making Skills self-documenting.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →