How Hermes Agent Integrates with Atropos for RL Training: Architecture & Implementation

Hermes Agent integrates with the RL training environment using Atropos through a specialized inheritance chain that bridges the Atropos RL server with Hermes' multi-turn tool-calling engine, enabling sandboxed rollouts with deterministic reward computation.

Hermes Agent ships with full Atropos integration that lets you train language models on multi-turn, tool-calling tasks. The NousResearch/hermes-agent repository provides an abstract base class that connects the Atropos rollout system to the Hermes tool-calling loop, allowing researchers to implement custom RL environments with access to terminal, file, and web tools. This integration supports both evaluation via OpenAI-compatible servers and full RL training with exact token IDs through VLLM ManagedServer.

Core Architecture: From Atropos to Hermes

The integration is built around a three-layer inheritance chain that maps Atropos primitives to Hermes capabilities:


atroposlib.BaseEnv  →  HermesAgentBaseEnv  →  ConcreteEnv

atroposlib.BaseEnv provides the generic RL-training server, worker pool, and WandB logging. HermesAgentBaseEnv (environments/hermes_base_env.py) adds the Hermes-specific plumbing: it configures the terminal sandbox backend (local, Docker, Modal), resolves toolsets for each rollout group, and runs the HermesAgentLoop inside an Atropos rollout. Concrete environments (e.g., TerminalTestEnv, HermesSweEnv, or user-defined subclasses) implement five abstract methods required by Atropos: setup, get_next_item, format_prompt, compute_reward, and evaluate.

The base class exposes a ToolContext (environments/tool_context.py) that gives reward functions direct access to the exact sandbox the model used during the rollout. This ensures deterministic verification—reward functions can invoke terminal, file, or web tools to check the model's work in the same environment that produced it.

Two-Phase Rollout System

Hermes Agent supports two operating modes depending on your training stage:

Phase 1 – OpenAI-compatible server targets evaluation, SFT data generation, and quick benchmarks. The server (OpenAI, OpenRouter, or VLLM) parses tool calls natively; Hermes sends tools= and receives structured tool_calls. Placeholder token IDs are generated for the Atropos pipeline, making this mode fast but unsuitable for RL algorithms requiring exact log-probs.

Phase 2 – VLLM ManagedServer enables full RL training with GRPO or PPO. The server returns raw text; Hermes uses a client-side tool-call parser (e.g., hermes, mistral, qwen) to reconstruct tool_calls from the model output. Real token IDs flow through Atropos, enabling precise advantage computation. The parser selection is controlled by the tool_call_parser configuration field.

Environment Configuration and Toolset Resolution

HermesAgentEnvConfig

HermesAgentBaseEnv expects a configuration subclass of HermesAgentEnvConfig (defined in environments/hermes_base_env.py). This config controls:

  • Toolset selection – enabled_toolsets, disabled_toolsets, or a probabilistic distribution resolved via toolset_distributions.sample_toolsets_from_distribution.
  • Terminal backend – terminal_backend (local, docker, modal). The selected backend is exported to environment variables (TERMINAL_ENV, TERMINAL_TIMEOUT, TERMINAL_LIFETIME) so all Hermes tools pick it up automatically.
  • Rollout limits – max_agent_turns, agent_temperature, tool_pool_size.
  • Phase-2 parser – tool_call_parser (default hermes).

Toolset Resolution

Before each group of rollouts, _resolve_tools_for_group() calls model_tools.get_tool_definitions() to generate the tool schema:

tools, valid_names = get_tool_definitions(
    enabled_toolsets=group_toolsets,
    disabled_toolsets=self.config.disabled_toolsets,
    quiet_mode=True,
)

This resolution happens once per group, not per trajectory, reducing overhead significantly. The returned schema list is passed to HermesAgentLoop so the LLM knows which tools are available for the current rollout.

Rollout Execution and Reward Computation

The collect_trajectory Method

collect_trajectory() (lines 5110-5230 in hermes_base_env.py) performs the core rollout logic:

  1. Create a unique task_id – guarantees sandbox isolation across parallel workers.
  2. Prepare the initial message list – combines the optional system prompt with the formatted user prompt from format_prompt(item).
  3. Choose the phase – _use_managed_server() decides whether to use a ManagedServer (Phase 2) or direct server (Phase 1). If Phase 2, it loads the selected parser from environments/tool_call_parsers.get_parser and enters a managed_server context.
  4. Run the agent loop – HermesAgentLoop(...).run(messages) (from environments/agent_loop.py) handles tool calls, executes them via a ThreadPoolExecutor(self.config.tool_pool_size) to avoid deadlocks with async back-ends, and returns an AgentResult.
  5. Compute reward – compute_reward(item, result, ctx) receives a fresh ToolContext(task_id) pointing at the exact sandbox used during the rollout. Reward functions can invoke any Hermes tool to verify the model's work.
  6. Package the trajectory – If the server produced real tokens (Phase 2), they are taken from result.managed_state["nodes"]; otherwise, placeholder tokens are generated from the conversation text.

Async-Safety Patches

Many Hermes tools (e.g., the Modal backend) spawn their own event loops via asyncio.run(). When called from inside Atropos' event loop, this raises RuntimeError: Cannot run the event loop while another loop is running. environments/patches.py (lines 44-86) replaces those calls with a dedicated background thread (_AsyncWorker) that runs its own loop, allowing safe execution during rollouts.

Logging and Visualization

add_rollouts_for_wandb() formats the full message trace into a compact, human-readable view via _format_trajectory_for_display. This appears in the WandB rollout table and includes tool calls, reasoning steps, and execution results, making debugging RL experiments straightforward.

Building a Custom Atropos Environment

Below is a minimal example that trains a model to generate Python functions and scores it by running the produced script in the same sandbox.


# my_env.py

from environments.hermes_base_env import HermesAgentBaseEnv, HermesAgentEnvConfig
from atroposlib.envs.server_handling.server_manager import APIServerConfig

class MyEnvConfig(HermesAgentEnvConfig):
    """Custom config – enable only terminal + file tools, use Modal sandbox."""
    pass  # all defaults are fine; you could set terminal_backend="modal" here

class MyEnv(HermesAgentBaseEnv):
    """Simple RL environment that asks the model to write a function."""
    name = "my-env"
    env_config_cls = MyEnvConfig

    @classmethod
    def config_init(cls):
        # Enable the terminal and file toolsets, run in Modal for isolation.

        env_cfg = MyEnvConfig(
            enabled_toolsets=["terminal", "file"],
            terminal_backend="modal",
            max_agent_turns=20,
            tool_pool_size=64,
        )
        # Use an OpenAI‑compatible server (OpenRouter in this example)

        server_cfg = [
            APIServerConfig(
                base_url="https://openrouter.ai/api/v1",
                model_name="anthropic/claude-sonnet-4.6",
                server_type="openai",
            )
        ]
        return env_cfg, server_cfg

    async def setup(self):
        # No external dataset – generate a trivial list of prompts.

        self.prompts = [
            "Write a function `add(a, b)` that returns a + b.",
            "Write a function `reverse(s)` that returns the reversed string.",
        ]
        self.idx = 0

    async def get_next_item(self):
        item = {"prompt": self.prompts[self.idx % len(self.prompts)]}
        self.idx += 1
        return item

    def format_prompt(self, item):
        # The model sees a plain user message.

        return item["prompt"]

    async def compute_reward(self, item, result, ctx):
        # Run the generated script in the same sandbox and check its output.

        # 1. Write the model's answer to a file.

        ctx.write_file("/workspace/solution.py", result.messages[-1]["content"])
        # 2. Execute it with the terminal tool.

        exec_res = ctx.terminal("python /workspace/solution.py")
        # 3. Reward is 1.0 if the script exits cleanly, otherwise 0.0.

        return 1.0 if exec_res["exit_code"] == 0 else 0.0

    async def evaluate(self, *_, **__):
        # Optional periodic evaluation – could run a few held‑out prompts.

        pass

if __name__ == "__main__":
    # CLI entry point – supports `serve`, `process`, `evaluate`

    MyEnv.cli()

Key implementation details:

  • ToolContext (ctx) gives direct access to the sandbox the model used via ctx.write_file() and ctx.terminal().
  • The terminal backend is set to Modal, which runs each rollout in an isolated cloud container.
  • No external dataset is required; the environment supplies its own prompts through setup() and get_next_item().

Running the Integration

Start the Atropos server and connect your environment:


# 1️⃣  Start the Atropos API (run in a separate terminal)

run-api

# 2️⃣  Run the environment in RL mode (connects to Atropos)

python my_env.py serve \
    --openai.model_name anthropic/claude-opus-4.6 \
    --env.max_agent_turns 30
  • Use process instead of serve to generate SFT trajectories without an Atropos server.
  • Use evaluate to run a benchmark-style evaluation loop.

Summary

  • HermesAgentBaseEnv (environments/hermes_base_env.py) bridges the Atropos RL server to the Hermes tool-calling loop via a three-layer inheritance chain.
  • ToolContext ensures reward functions can re-use the exact sandbox that produced the rollout, guaranteeing deterministic verification.
  • Phase 1 vs Phase 2 lets you choose between speed (OpenAI-compatible servers with placeholder tokens) and fidelity (VLLM ManagedServer with real token IDs and log-probs for GRPO/PPO).
  • Async patches in environments/patches.py enable safe use of Modal and Docker backends inside Atropos' event loop.
  • HermesAgentLoop (environments/agent_loop.py) provides the reusable multi-turn agent engine that executes tool calls during rollouts.

Frequently Asked Questions

How does Hermes Agent handle tool-calling during RL training with Atropos?

During Phase 2 training, Hermes Agent uses client-side parsers in environments/tool_call_parsers/ to reconstruct tool calls from raw model text generated by the VLLM ManagedServer. This allows Atropos to record exact token IDs and log-probs while Hermes still executes tools in the sandboxed environment. The HermesAgentLoop manages the multi-turn conversation, executing tools via a thread pool to prevent deadlocks with async backends.

What is the difference between Phase 1 and Phase 2 rollouts?

Phase 1 uses standard OpenAI-compatible servers (OpenAI, OpenRouter, or VLLM in chat mode) where the server parses tool calls natively. Hermes receives structured tool_calls directly, but Atropos only sees placeholder token IDs, making this suitable for evaluation and SFT data generation. Phase 2 uses VLLM ManagedServer returning raw text; Hermes parses tool calls client-side, allowing Atropos to capture real token IDs and compute precise advantages for RL algorithms like GRPO or PPO.

How can reward functions access the sandbox environment?

Reward functions receive a ToolContext instance (from environments/tool_context.py) via the compute_reward(item, result, ctx) method. This context provides methods like ctx.terminal(), ctx.write_file(), and ctx.read_file() that point to the exact sandbox container (local, Docker, or Modal) used during the rollout. This design ensures that verification happens in the same environment where the model executed its tools, eliminating inconsistencies.

Which files should I modify to create a custom RL environment?

Start by subclassing HermesAgentBaseEnv in environments/hermes_base_env.py and implementing the five required methods: setup, get_next_item, format_prompt, compute_reward, and evaluate. Define your configuration in a subclass of HermesAgentEnvConfig to control toolsets and terminal backends. If you need custom tool-call parsing for Phase 2 training, add a parser to environments/tool_call_parsers/. Refer to website/docs/developer-guide/environments.md for the full architecture specification.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →