# Dynamic Tool Discovery: How to Reduce Agent Context Overhead in LLM Systems

> Learn dynamic tool discovery to reduce LLM agent context overhead. Request tools on-demand and avoid token bloat. Optimize your AI agent performance now.

- Repository: [Bojie Li/ai-agent-book](https://github.com/bojieli/ai-agent-book)
- Tags: deep-dive
- Published: 2026-08-18

---

**Dynamic tool discovery reduces agent context overhead by letting LLMs request only the tools they need on-demand, avoiding the token bloat of pre-loading every available tool schema.**

Traditional LLM agents embed all possible tool definitions in their initial system prompt, forcing the model to process irrelevant metadata and pay inflated token costs. The `bojieli/ai-agent-book` repository demonstrates an **active discovery** pattern that keeps the initial prompt minimal, retrieves tools iteratively, and quantifies the savings. This approach is implemented in the `chapter4/active-tool-selection/` module, which provides working code you can adapt to your own agent architecture.

---

## The Problem: Token Bloat from Passive Tool Injection

When an agent needs to call external APIs—GitHub, filesystem operations, web search, databases—the straightforward approach loads every tool schema upfront. In [`agent.py`](https://github.com/bojieli/ai-agent-book/blob/main/agent.py), the `PassiveToolAgent` class demonstrates this baseline: it concatenates all `ToolDefinition` objects into the system message before the first LLM call.

The cost grows linearly with your tool catalog. A production system with 50+ tools can easily add thousands of tokens to every request, with most going unused for any specific task. This impairs the model's reasoning capacity and increases API costs predictably.

---

## The Solution: Active Tool Discovery Architecture

The repository implements three agent variants in [`chapter4/active-tool-selection/agent.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter4/active-tool-selection/agent.py):

| Agent Type | Strategy | Use Case |
|------------|----------|----------|
| **PassiveToolAgent** | Pre-loads all tools | Baseline for comparison |
| **RetrievalToolAgent** | Single-shot top-k retrieval | Fixed budget, predictable latency |
| **ActiveToolAgent** | Dynamic on-demand discovery | Maximum efficiency, iterative tasks |

`ActiveToolAgent` achieves the lowest context overhead through a tight feedback loop between the LLM and a **semantic router**.

---

## Core Components of the Discovery System

### Tool Knowledge Base

The catalog is defined in [`chapter4/active-tool-selection/tool_knowledge_base.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter4/active-tool-selection/tool_knowledge_base.py). It organizes capabilities hierarchically:

- `ServerDefinition` objects represent domains (GitHub, filesystem, web, email)
- Each server bundles `ToolDefinition` objects with metadata, descriptions, and JSON schemas

Example tools include `github_search_repos`, `fs_read_file`, and `web_get`. The structure supports clean separation between tool metadata and runtime execution.

### Semantic Router with Hierarchical TF-IDF

The `SemanticRouter` class in [`chapter4/active-tool-selection/semantic_router.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter4/active-tool-selection/semantic_router.py) performs two-stage retrieval:

1. **Server selection**: Ranks server descriptions using a pre-built TF-IDF index (`_build_server_index`)
2. **Tool selection**: Within the top server, ranks tools using per-server indices (`_build_tool_indices`)

Both stages use cosine similarity against a query string formed from the model's natural-language request. The `route_request` method returns the best-matching `{server, tool}` pair.

This hierarchical approach is more accurate than flat retrieval because server context disambiguates tool names. "Search" means something different for GitHub versus email.

### Structured Request Parser

The `StructuredRequestParser` class (same file) handles the LLM-to-router interface. It extracts `<tool_request>` blocks from model output:

```python
from chapter4.active-tool-selection.semantic_router import StructuredRequestParser

request = StructuredRequestParser.format_request(
    server_desc="GitHub repository platform",
    tool_desc="search for repositories with a specific keyword"
)
print(request)

```

Output:

```

<tool_request>
server: GitHub repository platform
tool: search for repositories with a specific keyword
</tool_request>

```

The parser's `parse_request` method converts this back to a machine-readable dict for the router.

---

## How ActiveToolAgent Reduces Context Overhead

The discovery loop in `ActiveToolAgent.execute_task` (lines 52-84 of [`agent.py`](https://github.com/bojieli/ai-agent-book/blob/main/agent.py)) follows this flow:

1. **Minimal system prompt** — Contains only the instruction that tools are available on request, no schemas
2. **Task execution** — LLM processes the user task with current context
3. **Gap detection** — If capabilities are missing, model emits `<tool_request>`
4. **Semantic routing** — `_handle_tool_request` calls `SemanticRouter.route_request` to find relevant tools
5. **Surgical injection** — Only discovered tool schemas (via `ToolDefinition.to_schema`) are appended
6. **Continued execution** — Next LLM call includes new tools via `tool_choice="auto"`
7. **Bounded iteration** — Repeats up to `MAX_TOOL_REQUESTS` (configured in [`config.py`](https://github.com/bojieli/ai-agent-book/blob/main/config.py))

The agent tracks `tokens_used`, `tool_requests`, and `tools_loaded` in its `metrics` dictionary, enabling empirical comparison with passive approaches.

---

## Practical Code Examples

### Basic Active Discovery Usage

```python
from chapter4.active-tool-selection.agent import ActiveToolAgent

agent = ActiveToolAgent()
result = agent.execute_task(
    "Summarize the most starred Python repositories on GitHub that mention 'machine learning' in their README."
)

print("Final response:", result["response"])
print("Metrics:", result["metrics"])
print("Tools loaded:", result["tools_loaded"])

```

The model first recognizes it needs GitHub search capabilities, requests `github_search_repos`, receives the schema, calls the tool (simulated in demo), and produces the final summary. Only one tool schema was ever injected.

### Comparing Token Efficiency

```python
from chapter4.active-tool-selection.agent import ActiveToolAgent, PassiveToolAgent

task = "List all `.py` files in the current directory and count how many contain the word 'async'."

active = ActiveToolAgent()
passive = PassiveToolAgent()

active_res = active.execute_task(task)
passive_res = passive.execute_task(task)

print("Active tokens used:", active_res["metrics"]["tokens_used"])
print("Passive tokens used:", passive_res["metrics"]["tokens_used"])

```

The active agent loads only `fs_list_directory` and `fs_search_files` on demand. The passive agent pre-loads 50+ tools. The savings scale with catalog size and task specificity.

---

## Configuration and Tuning

Behavior is controlled via [`chapter4/active-tool-selection/config.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter4/active-tool-selection/config.py):

- `MAX_TOOL_REQUESTS` — Hard limit on discovery iterations per task
- Similarity thresholds for the semantic router
- Model selection and temperature

Adjust these based on your latency budget and confidence in the router's precision.

---

## Summary

- **Dynamic tool discovery** eliminates upfront schema injection, keeping initial prompts minimal
- **Hierarchical TF-IDF routing** in `SemanticRouter` matches natural-language requests to precise tool definitions
- **Structured request protocol** (`<tool_request>` blocks) lets the LLM signal capability gaps without parser complexity
- **Quantifiable metrics** in `ActiveToolAgent` enable systematic optimization of the token/context tradeoff
- The complete implementation in `bojieli/ai-agent-book` provides production-ready patterns for agent architects

---

## Frequently Asked Questions

### How much token reduction can dynamic tool discovery achieve?

The savings depend on your tool catalog size and task specificity. In the repository's benchmark ([`demo_comparison.py`](https://github.com/bojieli/ai-agent-book/blob/main/demo_comparison.py)), active discovery typically reduces tool-related tokens by 80-95% compared to passive injection when only 1-3 tools are needed from a catalog of 50+. The `metrics["tokens_used"]` field tracks this directly for your own workloads.

### What happens if the semantic router selects the wrong tool?

The LLM can issue another `<tool_request>` in the next turn, up to `MAX_TOOL_REQUESTS`. The conversational context now includes the previously discovered (but incorrect) tool, which helps disambiguate the follow-up request. In practice, the two-stage server-then-tool routing in `SemanticRouter.route_request` achieves high precision on typical queries.

### Is dynamic discovery slower than passive injection?

Yes, by one round-trip per tool discovery. However, reduced context size improves per-request latency and cost, often net-neutral or favorable for multi-turn tasks. For latency-critical single-turn tasks, use `RetrievalToolAgent` which performs one retrieval step upfront without the conversational loop.

### Can I adapt this pattern to my own tool registry?

Absolutely. Implement the `ServerDefinition` and `ToolDefinition` interfaces from [`tool_knowledge_base.py`](https://github.com/bojieli/ai-agent-book/blob/main/tool_knowledge_base.py) for your capabilities. The `SemanticRouter` indexes any description strings you provide. Ensure your agent's system prompt instructs the model to emit `<tool_request>` blocks when capabilities are missing, matching the format expected by `StructuredRequestParser`.