How to Design Agent Skill Orchestration and Active Tool Discovery
Agent skill orchestration guides LLMs through predefined state-machine workflows with role-based tool permissions, while active tool discovery enables runtime expansion of capabilities via semantic search when the current toolset is insufficient.
The bojieli/ai-agent-book repository demonstrates production-ready patterns for building autonomous agents that can follow structured multi-role workflows and dynamically discover missing capabilities. This implementation combines rigid state-machine orchestration with flexible runtime tool discovery to create systems that behave like human teams—delegating tasks to specialists and requesting expertise on demand.
Skill Orchestration: State-Machine Workflows with Tool Gating
Skill orchestration treats agent execution as a directed workflow through predefined roles (triage → research → coding → data_analysis → writing). Each skill operates within a sandboxed environment with a fixed system prompt and an explicit whitelist of authorized tools.
Core Components
The orchestration layer resides in three primary files:
skill_orchestrator.py– Contains theSkillOrchestratorclass that manages state transitions, enforces permission boundaries, and aggregates performance metrics.roles.py– Defines skill specifications including system prompts and theSKILL_TOOLSmapping that authorizes specific tool sets per role.tools.py– Hosts the global tool registry (TOOL_IMPLEMENTATIONSandTOOL_SCHEMAS) that serves as the source of truth for available capabilities.
How Skill Orchestration Works
The SkillOrchestrator class implements a five-stage enforcement mechanism:
- System Prompt Construction – A static
SKILL_SYSTEM_PROMPTis built once during initialization and attached to every LLM request via_messages_for_api. - Skill Loading via Tool Call – The model must explicitly invoke
load_skill(name="triage")to begin. The orchestrator validates the request, loads the markdown documentation fromSKILL_ROOT, and caches it in_skill_cache. - Tool Gating – After loading, only tools listed in
SKILL_TOOLS[skill]are exposed to the model. Unauthorized tool calls return a policy-rejection message without executing. - History and Activity Logging – Every assistant message, tool call, and final response is appended to
self.history, while a lightweightself.activitylog tracks which role performed which action for audit trails. - Metrics Aggregation – The
summary()method computes token usage, cache-hit statistics, and step counts, producing reproducible performance reports.
from openai import OpenAI
from chapter10.multi_role_transfer.skill_orchestrator import SkillOrchestrator
client = OpenAI()
agent = SkillOrchestrator(
client=client,
model="gpt-5.6-luna",
max_steps=15,
verbose=False,
)
# The orchestrator automatically enforces skill-loading order and tool permissions
final_answer = agent.run("请帮我分析 2023 年苹果公司的财务报表并写一段结论。")
print(final_answer)
print(agent.summary()) # token usage, cache hits, step counts
Active Tool Discovery: Runtime Capability Expansion
Active tool discovery allows an LLM to request missing specialist tools when the baseline catalog cannot satisfy a task requirement. The system queries a semantic index of all available tool schemas, injects the top-k matches into the conversation, and dynamically expands the usable tool set.
Experiment Infrastructure
The discovery mechanism relies on components from chapter4/active-tool-discovery/:
run_exact_experiment.py– The experiment driver that initializes the baseline catalog, manages the discovery loop, and verifies final receipts.LocalEmbeddingIndex– A dense-vector index built from all MCP tool schemas that enables semantic similarity search.DISCOVER_SCHEMA– The JSON schema definition for thediscover_toolsfunction that the model can invoke.- Validation Helpers –
parse_action,mcp_receipt, andgrade_planhandle action validation, execution logging, and accuracy computation.
The Discovery Workflow
The active discovery process follows six distinct phases:
- Baseline Initialization – The experiment begins with a restricted catalog (
BASE_TOOL_NAMES) containing only generic utilities. - Model Turn with Status – The LLM receives a system prompt describing the task and current available tools, displayed in a status bar.
- Discovery Request – When the model selects
action="discover_tools", the orchestrator queriesLocalEmbeddingIndex.search()to retrieve the five most semantically similar schemas for the described need. - Schema Injection – Discovered schemas are rendered as JSON, hashed, and appended to the conversation history as user messages. The
availabletool set updates to include the new schema names. - Specialist Tool Execution – The model may now call newly exposed tools (e.g.,
arxiv_search,github_list_contributors). The runtime executes real MCP calls, records receipts viamcp_receipt, and verifies observations are substantive. - Verification and Acceptance –
derive_acceptancechecks that every required capability slot has a valid receipt, validates token counts, and confirms the campaign meets published acceptance gates.
import asyncio
from chapter4.active_tool_discovery.run_exact_experiment import run
# Run control vs treatment experiments with dynamic tool discovery
records = asyncio.run(run(campaign_id="demo-2024"))
summary = safe_summary(records) # aggregates accuracy & completion stats
print(summary["treatment"]["mean_tool_selection_accuracy"])
Shared Architectural Principles
Both patterns leverage common infrastructure design decisions that ensure reliability and auditability.
Shared Conversation History
All roles and tools operate on the same history list, ensuring every subsequent step sees the full conversational context without information silos.
Explicit Tool Schema Exposure
The model only receives JSON schemas for tools explicitly authorized for the current skill or discovered via the semantic index. This prevents accidental leakage of privileged capabilities not relevant to the active task.
Deterministic Auditing
Every LLM call, tool invocation, and role transfer is logged with timestamps, token counts, and sha256_bytes hash-chains. This enables reproducible verification through the derive_acceptance protocol.
Graceful Failure Handling
When policy violations occur—such as calling an unauthorized tool or attempting invalid role transfers—the orchestrator returns structured error messages and logs the event. This allows the LLM to recover without crashing the pipeline.
Summary
- Skill orchestration enforces deterministic workflows through
SkillOrchestrator, usingload_skill()calls andSKILL_TOOLSwhitelists to gate tool access by role. - Active tool discovery enables autonomous capability expansion via
LocalEmbeddingIndexsemantic search, injecting discovered schemas into the conversation history at runtime. - Both patterns share unified conversation history, explicit schema exposure, and deterministic auditing with hash-chained receipts.
- The implementation in
bojieli/ai-agent-bookprovides production-ready examples inskill_orchestrator.pyandrun_exact_experiment.py.
Frequently Asked Questions
What is the difference between skill orchestration and active tool discovery?
Skill orchestration manages deterministic, state-machine workflows where the agent transitions through predefined roles (triage, research, coding) with fixed tool permissions. Active tool discovery is a dynamic mechanism that allows the agent to request and integrate new tools at runtime when the current catalog proves insufficient for a specific task.
How does the orchestrator prevent unauthorized tool usage?
The SkillOrchestrator maintains a SKILL_TOOLS mapping that defines authorized tools per skill. After a skill is loaded via load_skill(), the orchestrator filters the global TOOL_SCHEMAS registry, exposing only whitelisted tools to the model. Attempts to call unauthorized tools return policy-rejection messages without execution.
What happens when the LLM requests a tool that isn't in the current catalog?
When the model invokes action="discover_tools", the system queries the LocalEmbeddingIndex to perform semantic similarity search against all available MCP schemas. The top-k matches are injected into the conversation history as JSON schemas, and the available tool set updates dynamically, allowing immediate use of the discovered tools in subsequent turns.
How does the system ensure reproducibility and auditing?
Every operation generates structured logs with timestamps, token counts, and sha256_bytes hashes. The derive_acceptance function verifies that all required capability slots have valid receipts, checks that dynamic token counts match expectations, and validates the campaign against published acceptance gates, creating a cryptographically verifiable audit trail.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →