Dynamic Tool Discovery: How to Reduce Agent Context Overhead in LLM Systems

Dynamic tool discovery reduces agent context overhead by letting LLMs request only the tools they need on-demand, avoiding the token bloat of pre-loading every available tool schema.

Traditional LLM agents embed all possible tool definitions in their initial system prompt, forcing the model to process irrelevant metadata and pay inflated token costs. The bojieli/ai-agent-book repository demonstrates an active discovery pattern that keeps the initial prompt minimal, retrieves tools iteratively, and quantifies the savings. This approach is implemented in the chapter4/active-tool-selection/ module, which provides working code you can adapt to your own agent architecture.


The Problem: Token Bloat from Passive Tool Injection

When an agent needs to call external APIs—GitHub, filesystem operations, web search, databases—the straightforward approach loads every tool schema upfront. In agent.py, the PassiveToolAgent class demonstrates this baseline: it concatenates all ToolDefinition objects into the system message before the first LLM call.

The cost grows linearly with your tool catalog. A production system with 50+ tools can easily add thousands of tokens to every request, with most going unused for any specific task. This impairs the model's reasoning capacity and increases API costs predictably.


The Solution: Active Tool Discovery Architecture

The repository implements three agent variants in chapter4/active-tool-selection/agent.py:

Agent Type Strategy Use Case
PassiveToolAgent Pre-loads all tools Baseline for comparison
RetrievalToolAgent Single-shot top-k retrieval Fixed budget, predictable latency
ActiveToolAgent Dynamic on-demand discovery Maximum efficiency, iterative tasks

ActiveToolAgent achieves the lowest context overhead through a tight feedback loop between the LLM and a semantic router.


Core Components of the Discovery System

Tool Knowledge Base

The catalog is defined in chapter4/active-tool-selection/tool_knowledge_base.py. It organizes capabilities hierarchically:

  • ServerDefinition objects represent domains (GitHub, filesystem, web, email)
  • Each server bundles ToolDefinition objects with metadata, descriptions, and JSON schemas

Example tools include github_search_repos, fs_read_file, and web_get. The structure supports clean separation between tool metadata and runtime execution.

Semantic Router with Hierarchical TF-IDF

The SemanticRouter class in chapter4/active-tool-selection/semantic_router.py performs two-stage retrieval:

  1. Server selection: Ranks server descriptions using a pre-built TF-IDF index (_build_server_index)
  2. Tool selection: Within the top server, ranks tools using per-server indices (_build_tool_indices)

Both stages use cosine similarity against a query string formed from the model's natural-language request. The route_request method returns the best-matching {server, tool} pair.

This hierarchical approach is more accurate than flat retrieval because server context disambiguates tool names. "Search" means something different for GitHub versus email.

Structured Request Parser

The StructuredRequestParser class (same file) handles the LLM-to-router interface. It extracts <tool_request> blocks from model output:

from chapter4.active-tool-selection.semantic_router import StructuredRequestParser

request = StructuredRequestParser.format_request(
    server_desc="GitHub repository platform",
    tool_desc="search for repositories with a specific keyword"
)
print(request)

Output:


<tool_request>
server: GitHub repository platform
tool: search for repositories with a specific keyword
</tool_request>

The parser's parse_request method converts this back to a machine-readable dict for the router.


How ActiveToolAgent Reduces Context Overhead

The discovery loop in ActiveToolAgent.execute_task (lines 52-84 of agent.py) follows this flow:

  1. Minimal system prompt — Contains only the instruction that tools are available on request, no schemas
  2. Task execution — LLM processes the user task with current context
  3. Gap detection — If capabilities are missing, model emits <tool_request>
  4. Semantic routing — _handle_tool_request calls SemanticRouter.route_request to find relevant tools
  5. Surgical injection — Only discovered tool schemas (via ToolDefinition.to_schema) are appended
  6. Continued execution — Next LLM call includes new tools via tool_choice="auto"
  7. Bounded iteration — Repeats up to MAX_TOOL_REQUESTS (configured in config.py)

The agent tracks tokens_used, tool_requests, and tools_loaded in its metrics dictionary, enabling empirical comparison with passive approaches.


Practical Code Examples

Basic Active Discovery Usage

from chapter4.active-tool-selection.agent import ActiveToolAgent

agent = ActiveToolAgent()
result = agent.execute_task(
    "Summarize the most starred Python repositories on GitHub that mention 'machine learning' in their README."
)

print("Final response:", result["response"])
print("Metrics:", result["metrics"])
print("Tools loaded:", result["tools_loaded"])

The model first recognizes it needs GitHub search capabilities, requests github_search_repos, receives the schema, calls the tool (simulated in demo), and produces the final summary. Only one tool schema was ever injected.

Comparing Token Efficiency

from chapter4.active-tool-selection.agent import ActiveToolAgent, PassiveToolAgent

task = "List all `.py` files in the current directory and count how many contain the word 'async'."

active = ActiveToolAgent()
passive = PassiveToolAgent()

active_res = active.execute_task(task)
passive_res = passive.execute_task(task)

print("Active tokens used:", active_res["metrics"]["tokens_used"])
print("Passive tokens used:", passive_res["metrics"]["tokens_used"])

The active agent loads only fs_list_directory and fs_search_files on demand. The passive agent pre-loads 50+ tools. The savings scale with catalog size and task specificity.


Configuration and Tuning

Behavior is controlled via chapter4/active-tool-selection/config.py:

  • MAX_TOOL_REQUESTS — Hard limit on discovery iterations per task
  • Similarity thresholds for the semantic router
  • Model selection and temperature

Adjust these based on your latency budget and confidence in the router's precision.


Summary

  • Dynamic tool discovery eliminates upfront schema injection, keeping initial prompts minimal
  • Hierarchical TF-IDF routing in SemanticRouter matches natural-language requests to precise tool definitions
  • Structured request protocol (<tool_request> blocks) lets the LLM signal capability gaps without parser complexity
  • Quantifiable metrics in ActiveToolAgent enable systematic optimization of the token/context tradeoff
  • The complete implementation in bojieli/ai-agent-book provides production-ready patterns for agent architects

Frequently Asked Questions

How much token reduction can dynamic tool discovery achieve?

The savings depend on your tool catalog size and task specificity. In the repository's benchmark (demo_comparison.py), active discovery typically reduces tool-related tokens by 80-95% compared to passive injection when only 1-3 tools are needed from a catalog of 50+. The metrics["tokens_used"] field tracks this directly for your own workloads.

What happens if the semantic router selects the wrong tool?

The LLM can issue another <tool_request> in the next turn, up to MAX_TOOL_REQUESTS. The conversational context now includes the previously discovered (but incorrect) tool, which helps disambiguate the follow-up request. In practice, the two-stage server-then-tool routing in SemanticRouter.route_request achieves high precision on typical queries.

Is dynamic discovery slower than passive injection?

Yes, by one round-trip per tool discovery. However, reduced context size improves per-request latency and cost, often net-neutral or favorable for multi-turn tasks. For latency-critical single-turn tasks, use RetrievalToolAgent which performs one retrieval step upfront without the conversational loop.

Can I adapt this pattern to my own tool registry?

Absolutely. Implement the ServerDefinition and ToolDefinition interfaces from tool_knowledge_base.py for your capabilities. The SemanticRouter indexes any description strings you provide. Ensure your agent's system prompt instructs the model to emit <tool_request> blocks when capabilities are missing, matching the format expected by StructuredRequestParser.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →