What is the Skills-over-MCP Pattern for AI Agents?
The Skills-over-MCP pattern is an architectural strategy that combines the Model Context Protocol (MCP) for universal tool interoperability with a progressive-disclosure mechanism called Skills to keep prompts compact while giving agents on-demand access to thousands of capabilities.
The Skills-over-MCP pattern addresses the fundamental tension between agent flexibility and context window limitations in modern AI systems. As documented in the bojieli/ai-agent-book repository, this approach leverages standardized protocol interoperability alongside intelligent capability discovery to solve both framework fragmentation and the "choice overload" problem. By decoupling tool availability from immediate prompt inclusion, agents can maintain lean context windows while retaining access to expansive tool ecosystems.
Core Components of the Pattern
The architecture rests on two complementary mechanisms that work together to provide scalable agent capabilities.
MCP (Model Context Protocol)
MCP is a standardized, client-server protocol for exposing tools, read-only resources, and prompt templates to LLMs. According to the source analysis of book-en/chapter4.md, MCP solves the interoperability problem by decoupling tool definitions from specific agent frameworks. This standardization allows the same tool to function across Cursor, Claude Desktop, OpenClaw, and other MCP-compatible environments without framework-specific rewrites.
Skills (Progressive Disclosure)
Skills provide a progressive-disclosure mechanism that stores only a thin catalog—containing just a name and description—within the agent's prompt. The full instructions, scripts, or documentation for a capability are retrieved on-demand when the model identifies a specific capability gap. As noted in the repository's test files (tests/test_ch4_semantic_router_stop_words_query.py, lines 140-148), this approach prevents token blow-up when hundreds of potential tools exist, addressing the choice-overload problem that plagues flat tool registries.
How the Skills-over-MCP Architecture Works
The pattern follows a six-stage execution flow that preserves cache efficiency while enabling dynamic capability expansion:
-
Static Prefix Injection – At session start, the prompt contains only a skill index: a minimal catalog listing each Skill's identifier and brief description. No full tool schemas are present initially.
-
Capability Gap Detection – During reasoning, the model recognizes it lacks a needed operation (e.g., "I need to fetch live stock prices").
-
Skill Lookup – The model issues a request to a Skill loader, typically implemented as a
discover_toolsmeta-tool. -
MCP-Backed Retrieval – The Skill loader uses the MCP protocol to query an MCP server for matching tool definitions or documentation. The server returns the complete schema or implementation script.
-
Dynamic Injection – The returned definition is appended to the conversation history after the static prefix. Because this content is placed at the end of the context, the KV-Cache for the prefix remains intact, limiting expensive cache invalidation.
-
Execution – The agent now possesses a fully-specified tool (via MCP) or concrete script (via the Skill) and can invoke it through standard protocol calls.
+-------------------+ +-------------------+ +-------------------+
| Agent Prompt | static | Skill Index | on-demand| Skill Loader |
| (thin catalog) |--------->| (name+desc) |--------->| (MCP client) |
+-------------------+ +-------------------+ +-------------------+
|
v
+-------------------+
| MCP Server |
| (tool schemas, |
| resources, prompts)|
+-------------------+
The flow moves from a lightweight prompt through the Skill loader to the MCP server, returning full definitions only when necessary.
Practical Implementation Guide
Defining the Static Skill Index
The agent receives a minimal catalog as part of its system prompt, eliminating the need to load full schemas upfront:
skill_index = """
Available Skills:
- name: pdf_summarizer
description: Summarise a PDF document.
- name: stock_price
description: Retrieve the latest stock price for a ticker.
- name: image_caption
description: Generate a caption for an image.
"""
Building the Dynamic Skill Loader
The Skill loader functions as an MCP client that retrieves full tool specifications on demand:
import json, requests
MCP_ENDPOINT = "http://localhost:8000" # MCP server address
def discover_skill(requested_capability: str) -> dict:
"""
Ask the MCP server for a tool that matches the requested capability.
Returns the full tool schema (which the model can later inject).
"""
payload = {
"type": "tool/search",
"query": requested_capability,
"limit": 3
}
resp = requests.post(f"{MCP_ENDPOINT}/tools/search", json=payload)
resp.raise_for_status()
return resp.json() # e.g. {"tools": [{...tool schema...}]}
Agent Execution Loop
The agent detects missing capabilities and triggers the discovery workflow:
def agent_loop(user_input: str):
# 1. Agent tries to solve the task with existing knowledge.
if "stock price" in user_input.lower():
# 2. Detect missing capability → ask the loader.
result = discover_skill("stock price")
# 3. Inject the returned tool schema into the conversation.
inject_into_context(json.dumps(result, indent=2))
# 4. Now the model can call the tool via MCP.
call_mcp_tool("stock_price", {"ticker": "AAPL"})
else:
# Handle other cases...
pass
The inject_into_context call appends the tool schema to the conversation after the static prefix, preserving the KV-Cache for the initial context.
MCP Tool Schema Example
Once discovered, the full schema enables precise tool invocation:
{
"name": "stock_price",
"description": "Fetch the latest price for a given ticker symbol.",
"parameters": {
"ticker": {"type": "string", "description": "Ticker symbol, e.g. AAPL"}
},
"returns": {"type": "object", "properties": {"price": {"type": "number"}}}
}
Key Source Files and References
The bojieli/ai-agent-book repository provides both theoretical foundations and concrete implementations:
-
book-en/chapter4.md– Contains the primary narrative describing the "From MCP to Skills" section, explaining how the pattern solves the problem of too many tools. -
slides/lesson-08.md(lines 20-24) – Provides visual explanations of progressive disclosure and on-demand capability loading, illustrating the conceptual flow between components. -
chapter2/agent-skills-ppt/skills/pptx/SKILL.md– Demonstrates a concrete Skill implementation following the progressive disclosure model, showing how complex capabilities (PowerPoint manipulation) are packaged as loadable units. -
chapter9/ai-style-skill/skill/SKILL.md– Exhibits how Skills can be expressed as markdown files that agents load dynamically, including full implementation details that remain absent from the static prompt until requested. -
tests/test_ch4_semantic_router_stop_words_query.py(lines 140-148) – Contains test implementations validating the semantic routing logic underlying the Skills-over-MCP discovery mechanism.
Summary
- Skills-over-MCP combines MCP's universal interoperability with a progressive-disclosure layer to prevent token overflow.
- The static prompt contains only lightweight skill descriptors (name + description), keeping initial context small.
- Full tool schemas and implementations are retrieved on-demand via standard MCP protocol calls.
- Dynamic injection occurs after the static prefix, preserving KV-Cache efficiency and minimizing recomputation costs.
- This architecture enables agents to safely leverage thousands of tools without suffering from choice overload or context window exhaustion.
Frequently Asked Questions
How does Skills-over-MCP differ from standard MCP implementations?
Standard MCP implementations typically expose all available tools to the agent at startup, which leads to choice overload and excessive token consumption when tool counts grow large. Skills-over-MCP introduces a discovery layer that keeps the prompt minimal—containing only skill names and descriptions—while loading full tool schemas dynamically via the discover_tools mechanism only when the agent identifies a specific need.
Can existing MCP servers be used with the Skills-over-MCP pattern?
Yes, the pattern is fully backward-compatible with existing MCP servers. The Skill loader acts as a standard MCP client that queries existing endpoints (such as /tools/search) to retrieve tool definitions dynamically. No modifications are required to the MCP server itself; the innovation occurs on the agent side through the progressive disclosure mechanism and the Skill loader's cache-aware injection strategy.
What prevents the context window from growing indefinitely with dynamic injection?
The architecture relies on progressive disclosure rather than cumulative accumulation. While the retrieved skill definition temporarily increases context size for the current task, implementations typically manage conversation history to remove or compress stale tool definitions once the immediate task completes. Additionally, because definitions are appended after the static prefix, the KV-Cache for the prefix remains stable, minimizing computational overhead even as dynamic content cycles in and out.
Where can I find the official specification for this pattern?
The primary documentation resides in Chapter 4 of the bojieli/ai-agent-book repository, specifically the section titled "From MCP to Skills: Solving the problem of too many tools." The Model Context Protocol specification and Skills-over-MCP working group materials are also referenced in the chapter's footnotes, providing canonical definitions of the protocol and architectural patterns.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →