How Code Is Treated as a Tool for Creating New Tools in AI Agents

In the bojieli/ai-agent-book repository, a self-evolving AI agent treats validated Python code as a first-class tool by sandboxing execution, serializing functions to JSON, and exposing a searchable library interface that enables dynamic discovery and reuse without reimplementation.

This architecture demonstrates how modern AI agents can expand their own capabilities by treating code as a tool for creating new tools. The implementation centers on the ToolLibrary class in chapter9/self-evolving-tools/tool_manager.py, which orchestrates a five-stage pipeline from initial code validation to persistent, reusable tool execution.

The Self-Evolving Tool Architecture

The system transforms arbitrary Python snippets into durable capabilities through a strict lifecycle: validation, packaging, persistence, discovery, and execution. Each stage is anchored in specific source files that enforce isolation and reproducibility.

Step 1: Validating Code in the Sandbox

Before any code becomes a reusable tool, the agent must verify it executes safely and correctly. The code_interpreter utility defined in chapter9/self-evolving-tools/base_tools.py (lines 42‑99) runs arbitrary Python snippets inside an isolated subprocess. This sandbox optionally installs third‑party packages into a local .sandbox_packages directory, capturing real execution output (or failure traces) that the model can trust before proceeding.

Step 2: Packaging Code as a Reusable Tool

Once validation succeeds, the agent invokes ToolLibrary.create_tool() in chapter9/self-evolving-tools/tool_manager.py (lines 50‑101). This method accepts five critical parameters: name, description, parameters (JSON Schema), code (the Python source), and optional test_args. The method performs a syntax check, optionally runs the code against the test arguments, and constructs a JSON record containing the tool’s metadata and full source code.

Step 3: Persisting Tools to Disk

The JSON record is written to the tool_library/ directory (referenced internally as LIBRARY_DIR). Specifically, line 95 in tool_manager.py executes (self.dir / f"{name}.json") to serialize each tool as a standalone file. This persistence layer ensures that tools survive process restarts and can be shared across different agent instances without requiring re‑execution of the original code generation logic.

Step 4: Discovering Existing Tools

On subsequent requests, the agent avoids redundant implementation by calling ToolLibrary.search_tools(query) (lines 104‑130 in tool_manager.py). This method scans all stored JSON records, performing substring matches against both the tool’s name and description fields. It returns ranked results, allowing the agent to discover that a suitable tool already exists for a given task.

Step 5: Executing Stored Tools

When the agent selects a stored tool, it calls ToolLibrary.execute_tool(name, arguments) (lines 155‑207). This method loads the stored JSON, injects the code into a temporary script, and invokes the internal _run_record() helper. The helper executes the run(**arguments) function inside the same sandbox environment used by code_interpreter, capturing results via a special marker (__TOOL_RESULT__) to ensure clean output isolation.

Practical Implementation: Creating and Using Custom Tools

The repository provides concrete patterns for interacting with the tool lifecycle. Below are runnable examples demonstrating the full flow from creation to execution.

Registering a New Tool with create_tool

To treat code as a reusable tool, define a Python function named run that accepts keyword arguments, then register it through the library interface:

from chapter9.self_evolving_tools.tool_manager import ToolLibrary

# Initialize the library (defaults to tool_library/ directory)

lib = ToolLibrary()

# Python code must define a run(**kwargs) function

my_code = """
def run(url: str):
    import requests
    r = requests.get(url, timeout=5)
    return {"status": r.status_code, "content": r.text[:200]}
"""

# Register the tool with validation

result = lib.create_tool(
    name="fetch_title",
    description="Fetch the first 200 characters of a web page.",
    parameters={"type": "object", "properties": {"url": {"type": "string"}}},
    code=my_code,
    test_args={"url": "https://example.com"}  # Optional validation run

)

print(result["message"])

# Output confirms creation after syntax and test-run checks

Result: The system saves the tool as tool_library/fetch_title.json only after the code passes both static syntax analysis and dynamic execution against the provided test_args.

Searching the Library with search_tools

Agents can query existing capabilities to avoid regenerating functionality:

hits = lib.search_tools("fetch web page")
print([t["name"] for t in hits["tools"]])

# → ['fetch_title']

This search operation scans the tool_library/ directory, comparing the query string against stored metadata without loading or executing the underlying code, making it lightweight for frequent lookups.

Invoking Stored Logic with execute_tool

When the agent determines that fetch_title solves the current task, it invokes the stored logic by name:

exec_res = lib.execute_tool(
    name="fetch_title",
    arguments={"url": "https://github.com/bojieli/ai-agent-book"}
)

print(exec_res)

# → {'success': True, 'result': {'status': 200, 'content': '<!doctype html>...'}}

The execute_tool method handles dependency resolution automatically. Because the sandbox maintains the .sandbox_packages directory used during initial validation, third‑party imports (like requests) remain available across executions.

Summary

  • Sandboxed validation via code_interpreter in base_tools.py ensures only working code enters the library.
  • ToolLibrary.create_tool() packages validated Python into JSON records containing metadata and source.
  • File-based persistence stores each tool as an individual JSON file in tool_library/, enabling survival across sessions.
  • search_tools() provides semantic discovery without code execution overhead.
  • execute_tool() reloads and runs stored code in an isolated environment using the run(**arguments) convention.
  • Code becomes a meta-tool: The agent uses existing tools (code_interpreter) to create new tools, demonstrating a self‑evolving capability loop.

Frequently Asked Questions

What makes the tool creation process "self-evolving"?

The process is self‑evolving because the agent uses its own existing tools—specifically code_interpreter—to validate and package new capabilities without human intervention. As implemented in bojieli/ai-agent-book, the agent can write Python code, test it in a sandbox, and immediately promote it to a first‑class tool via ToolLibrary.create_tool(), effectively expanding its own functional API through iteration.

Why does the system require a run(**kwargs) function signature?

The execute_tool() method in tool_manager.py (lines 155‑207) dynamically loads the stored source code and specifically invokes a function named run with the provided arguments dictionary unpacked as keyword arguments. This convention standardizes the interface between the library’s execution engine and arbitrary user‑generated code, ensuring that ToolLibrary can treat heterogeneous tools uniformly without parsing Abstract Syntax Trees at runtime.

How does the sandbox handle external dependencies like requests?

The code_interpreter tool in base_tools.py manages a local .sandbox_packages directory where third‑party packages are installed during the initial validation run. When execute_tool() later invokes _run_record(), it executes within the same subprocess environment, meaning imports resolved during the create_tool() phase remain available for all subsequent executions of that stored tool.

Where is the complete demonstration of this architecture?

The file chapter9/self-evolving-tools/demo.py contains a runnable end‑to‑end demonstration that chains web_search, code_interpreter, and ToolLibrary methods to show autonomous tool creation and reuse. Additionally, chapter9/self-evolving-tools/agent.py provides a reference implementation of an agent that coordinates these primitives to achieve self‑evolution in practice.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →