How the execute_code Tool Enables Programmatic Tool Calling with RPC in Hermes Agent

The execute_code tool generates a stub module that wraps Hermes tools in RPC calls over a Unix-domain socket, allowing LLM-generated Python scripts to programmatically invoke tools from inside a sandboxed child process.

The NousResearch Hermes Agent provides a secure code execution environment where large language models can write and run Python scripts that interact with external tools. The execute_code tool bridges this capability by establishing a Remote Procedure Call (RPC) layer over Unix-domain sockets, enabling programmatic tool calling from within sandboxed code while maintaining strict security boundaries.

RPC Architecture Overview

Dynamic Stub Module Generation

When execute_code is invoked, the parent process dynamically builds hermes_tools.py containing thin RPC wrappers for each sandbox-allowed tool such as web_search, read_file, and terminal. According to the source code in tools/code_execution_tool.py, the generate_hermes_tools_module function (lines 17-33) creates these wrappers, each calling a private _call function that serializes requests over a Unix-domain socket.

Unix-Domain Socket Server

The parent process opens a Unix-domain socket and starts the _rpc_server_loop in a background thread (lines 22-33 of tools/code_execution_tool.py). This server listens for JSON-encoded tool-call messages, validates each request against an allow-list, dispatches calls via the standard handle_function_call path, and returns results on the same socket.

Child Sandbox Initialization

The execute_code function creates a temporary directory and writes both the generated hermes_tools.py stub and the user-supplied script.py before launching a child Python process (lines 48-56). The socket path is exported to the child via the HERMES_RPC_SOCKET environment variable, establishing the communication channel.

RPC Request and Response Flow

Inside the sandbox, a call like web_search("cats") triggers the stub's _call method (lines 78-87), which opens the socket via _connect, sends a JSON payload structured as {"tool":"web_search","args":{...}}, and blocks until the parent replies. The stub unwraps the JSON response, returning native Python data structures to the executing script.

Security Controls and Resource Limits

The parent process enforces strict isolation through multiple mechanisms defined in execute_code (lines 60-73). Execution time is capped by DEFAULT_TIMEOUT, while stdout and stderr sizes are limited to prevent memory exhaustion. The total number of tool calls is restricted by DEFAULT_MAX_TOOL_CALLS. Dangerous environment variables are stripped before spawning the child, and the sandbox can only invoke tools present in the intersection of session-enabled tools and SANDBOX_ALLOWED_TOOLS.

Additionally, the build_execute_code_schema function (lines 31-44) dynamically constructs the OpenAI function-calling schema to list only currently available tools, preventing the model from attempting unavailable operations.

Integration with the Agent Tool Registry

The execute_code functionality integrates into the broader Hermes Agent architecture through several supporting files. In tools/registry.py, the tool is registered with the global tool registry using the dynamically built schema from build_execute_code_schema. The model_tools.py file handles integration specifics, including refunding an iteration for the tool's turn as noted in run_agent.py (lines 109-127). For UI display, gateway/run.py maps the execute_code tool to the snake emoji (🐍), providing visual identification in agent outputs.

Practical Examples of Programmatic Tool Calling

Multi-Step Web Search and Extraction

The following script demonstrates chaining multiple tool calls within a single sandbox execution:

from hermes_tools import web_search, web_extract, json_parse
import json, re

# Search for Python release notes

search = web_search("python 3.12 release notes", limit=3)
url = search["data"]["web"][0]["url"]

# Extract content from the first result

page = web_extract([url])["results"][0]["content"]

# Parse version number using regex

match = re.search(r'Python\s+3\.\d+\.\d+', page)
print(json.dumps({"found": match.group(0) if match else None}))

The parent aggregates the results after the child exits, returning a JSON payload like:

{
  "status": "success",
  "output": "{\"found\": \"Python 3.12.0\"}",
  "tool_calls_made": 2,
  "duration_seconds": 12.34
}

Safe Terminal Execution

Sandboxed scripts can invoke the terminal tool with safety constraints. The RPC server strips background process parameters and PTY options, ensuring only foreground execution:

from hermes_tools import terminal, json_parse
import json

listing = terminal("ls -1")
print(json.dumps({"files": json_parse(listing["output"])}))

Batch Processing with Control Flow

Programmatic calling supports loops and conditional logic, with each iteration generating separate RPC calls tracked by the parent:

from hermes_tools import web_search, write_file
import json

results = []
for i in range(5):
    resp = web_search(f"hermes agent release {i}", limit=1)
    url = resp["data"]["web"][0]["url"]
    results.append(url)

write_file("/tmp/urls.txt", "\n".join(results))
print(json.dumps({"saved": True, "count": len(results)}))

The parent tracks tool_calls_made across the loop to enforce DEFAULT_MAX_TOOL_CALLS.

Summary

  • The execute_code tool generates hermes_tools.py stubs that wrap permitted tools in RPC calls over Unix-domain sockets.
  • The _rpc_server_loop in the parent process validates and dispatches tool calls while the child sandbox runs untrusted code.
  • Communication occurs via HERMES_RPC_SOCKET, with the _call function handling JSON serialization and blocking I/O.
  • Security boundaries include timeouts, output limits, maximum tool call counts, and environment variable sanitization.
  • Results are aggregated in execute_code (lines 170-184) and returned as a single JSON payload, compressing multi-step workflows into one inference turn.

Frequently Asked Questions

How does the sandbox communicate with the parent process?

The child process connects to the parent via a Unix-domain socket whose path is exported through the HERMES_RPC_SOCKET environment variable. The generated stub module in hermes_tools.py uses this socket to send JSON-encoded tool requests and receive responses, as implemented in the _call and _connect functions within the stub header.

What happens if a script exceeds the tool call limit?

The parent process tracks the number of RPC calls via tool_calls_made and enforces DEFAULT_MAX_TOOL_CALLS (defined in tools/code_execution_tool.py lines 60-73). Once the limit is reached, subsequent tool calls from the sandbox will fail, and the final result payload will indicate the limit was exceeded while preserving any output generated before the restriction triggered.

Which tools are available inside the execute_code sandbox?

Only tools appearing in the intersection of session-enabled tools and the SANDBOX_ALLOWED_TOOLS list are accessible. The build_execute_code_schema function dynamically generates the OpenAI function schema to reflect only these permitted tools, ensuring the LLM cannot attempt to invoke restricted capabilities from within generated code.

How does the RPC layer handle errors?

The _rpc_server_loop validates incoming requests against an allow-list before dispatching via handle_function_call. Errors during tool execution are serialized as JSON responses and returned through the socket to the stub's _call function, which unwraps them as standard Python exceptions or error dictionaries that the sandboxed script can catch and handle programmatically.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →