How the Active Tool Discovery Experiment (4-5) Enables Agents to Proactively Select Tools
The active tool discovery experiment (4-5) enables agents to proactively select tools by starting with zero pre-loaded functions, allowing the LLM to declare capability gaps via structured XML requests, and then dynamically injecting only the requested tool schemas from a semantic knowledge base before the next inference cycle.
The bojieli/ai-agent-book repository demonstrates a paradigm shift in agent architecture through Chapter 4's active tool discovery experiment (4-5). Unlike traditional approaches that preload entire tool catalogs into the system prompt, this method lets the agent analyze its own needs and request specific capabilities on demand. By implementing this workflow in chapter4/active-tool-selection/agent.py, the experiment achieves a 3.2× speedup while reducing token usage by approximately 1,000 tokens per task compared to passive injection baselines.
Starting with Zero Tools: The Minimal Context Strategy
Traditional agents inject all 120+ available tool schemas at startup, consuming thousands of tokens before any user query arrives. The active discovery approach reverses this by initializing ActiveToolAgent with an empty tool set.
The System Prompt Design (agent.py L11-L19)
The agent's system message, generated by _create_system_message, explicitly instructs the model to request capabilities when it identifies a gap. Instead of listing available functions, the prompt describes a protocol:
def _create_system_message(self) -> str:
"""Create system message explaining active tool discovery."""
return """You are an autonomous AI agent with active tool discovery capabilities.
...
When you identify a capability gap, request tools using this format:
<tool_request>
server: ...
tool: ...
</tool_request>"""
This design pattern forces the LLM to perform introspection about its own capabilities before attempting to act.
Detecting Capability Gaps Through Structured Requests
After each model generation, the agent scans the output for explicit tool requests rather than immediate function calls.
Parsing the Tool Request Block (agent.py L88-L92)
The execute_task method invokes StructuredRequestParser.parse_request at line 89 to detect <tool_request> XML blocks in the model's response:
# Inside the execution loop
parsed_request = StructuredRequestParser.parse_request(response)
if parsed_request:
# Treat as capability gap - proceed to discovery
server, tool = parsed_request["server"], parsed_request["tool"]
If a request is present, the agent interrupts its normal execution flow to enter the discovery phase.
Semantic Routing to the Tool Knowledge Base
Once the agent identifies a needed capability, it must locate the exact tool definition among the 120+ available options without loading them all into context.
Request Construction and Routing (agent.py L75-L80)
The agent constructs a routing key by merging the server and tool identifiers, then queries the semantic knowledge base:
# Construct the lookup key
request_str = f"{server} {tool}"
# Route to the knowledge base
discovered_tools = self.router.route_request(request_str)
The SemanticRouter queries the tool knowledge base created by create_tool_knowledge_base in tool_knowledge_base.py (lines 48-64), which indexes tool definitions by domain (GitHub, filesystem, web, etc.). This returns only the most relevant ToolDefinition objects based on semantic similarity.
On-Demand Tool Injection and Execution
Discovery is worthless without schema injection. The agent dynamically modifies its available tool set for the subsequent LLM call.
Dynamic Schema Loading (agent.py L90-L95)
Discovered tools are appended to self.available_tools immediately upon retrieval:
# Add discovered tools to available set
for tool in discovered_tools:
if tool.name not in [t.name for t in self.available_tools]:
self.available_tools.append(tool)
LLM Request Construction (agent.py L45-L48)
The _call_llm method then includes only these dynamically loaded tools in the OpenAI API request's tools parameter. This ensures that the token budget is spent entirely on relevant capabilities rather than a bloated catalog of unused functions.
Iterative Refinement and the Active Loop
Complex tasks often require multiple waves of discovery. The execute_task method (lines 82-104) implements an iterative refinement loop:
- The agent receives a task
- It generates a response, potentially requesting tools
- If tools are requested, they are discovered and injected
- The loop repeats with the expanded context up to
config.MAX_TOOL_REQUESTStimes
This allows the agent to chain discoveries: realizing it needs a web download tool, then subsequently discovering it needs a parsing tool for the downloaded content, all within the same task execution.
Performance Results: 3.2× Speedup with ~1k Tokens
According to EXPERIMENT_LEDGER.md (lines 61-73), experiment 4-S demonstrated dramatic efficiency gains over passive injection:
- 3.2× speedup in task completion time
- ~1,000 tokens used per task versus the full catalog baseline
- Identical accuracy on task completion metrics
These results validate that proactive tool selection does not sacrifice capability while dramatically reducing latency and API costs.
Code Example: Running the Active Discovery Demo
You can observe this behavior by comparing ActiveToolAgent against PassiveToolAgent:
from chapter4.active-tool-selection.agent import ActiveToolAgent, PassiveToolAgent
# Active agent starts with no tools
active_agent = ActiveToolAgent()
task = "Download the latest README from the repository and count the number of lines."
active_result = active_agent.execute_task(task)
print("Tools loaded proactively:", active_result["tools_loaded"])
print("Token usage:", active_result["metrics"]["tokens_used"])
# Compare with passive injection baseline
passive_agent = PassiveToolAgent()
passive_result = passive_agent.execute_task(task)
print("Passive token usage:", passive_result["metrics"]["tokens_used"])
In the active flow, the model first emits a <tool_request> block for web_download, receives the schema, and then executes the function—never seeing the other 119 unused tool definitions.
Summary
- Minimal initialization:
ActiveToolAgentstarts with zero tool schemas, eliminating context pollution from thechapter4/active-tool-selection/agent.pyimplementation. - Explicit gap declaration: The LLM requests capabilities using structured
<tool_request>XML blocks parsed at line 89 viaStructuredRequestParser.parse_request. - Semantic retrieval: The
SemanticRouter.route_requestmethod (lines 78-80) queriestool_knowledge_base.pyto retrieve only relevantToolDefinitionobjects. - Dynamic injection: Discovered tools are appended to
self.available_tools(lines 90-95) and injected into the OpenAI request via_call_llm(lines 45-48). - Iterative capability building: The
execute_taskloop (lines 82-104) supports up toMAX_TOOL_REQUESTSdiscovery cycles per task. - Quantified efficiency: The approach achieves a 3.2× speedup and reduces token usage to ~1k tokens compared to passive baselines while maintaining full task accuracy.
Frequently Asked Questions
How does active tool discovery differ from standard RAG tool selection?
Active tool discovery differs from standard Retrieval-Augmented Generation (RAG) in that the LLM itself initiates the retrieval rather than the system guessing which tools the model might need. In the bojieli/ai-agent-book implementation, the agent emits a specific <tool_request> XML block when it identifies a capability gap, whereas RAG systems typically retrieve based on query similarity without the model's explicit introspection. This self-directed approach reduces false positives in tool retrieval.
What happens if the agent requests a tool that does not exist in the knowledge base?
If SemanticRouter.route_request cannot find a matching tool in the knowledge base created by create_tool_knowledge_base, it returns an empty list. According to the logic in agent.py lines 90-95, no tools are appended to self.available_tools, and the agent receives a system message indicating the request could not be fulfilled. The loop then continues to config.MAX_TOOL_REQUESTS, allowing the model to reformulate its request or proceed with available capabilities.
Can active tool discovery handle multi-step tasks requiring sequential tool use?
Yes, the iterative design in execute_task (lines 82-104) specifically supports multi-step discovery. The agent can request a web search tool in the first iteration, receive results that indicate a need for parsing, and then request a text extraction tool in the second iteration. Each cycle expands self.available_tools cumulatively until the task is complete or the maximum request limit is reached, enabling complex dependency chains without pre-loading all possible tools.
Does proactive tool selection reduce accuracy compared to having all tools available?
No. According to the EXPERIMENT_LEDGER.md entries (lines 61-73), the active discovery approach achieved identical task-completion accuracy compared to the passive injection baseline while using only the fraction of the token budget. The semantic routing in tool_knowledge_base.py ensures that relevant tools are retrieved with high precision, meaning the agent receives the capabilities it needs without the noise of irrelevant schemas that might confuse the model.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →