# How Hermes Agent Integrates with Atropos for RL Training: Architecture & Implementation

> Discover how Hermes Agent integrates with Atropos for RL training. Explore the architecture and implementation details for sandboxed rollouts and deterministic reward computation.

- Repository: [Nous Research/hermes-agent](https://github.com/NousResearch/hermes-agent)
- Tags: architecture
- Published: 2026-03-09

---

**Hermes Agent integrates with the RL training environment using Atropos through a specialized inheritance chain that bridges the Atropos RL server with Hermes' multi-turn tool-calling engine, enabling sandboxed rollouts with deterministic reward computation.**

Hermes Agent ships with full Atropos integration that lets you train language models on multi-turn, tool-calling tasks. The NousResearch/hermes-agent repository provides an abstract base class that connects the Atropos rollout system to the Hermes tool-calling loop, allowing researchers to implement custom RL environments with access to terminal, file, and web tools. This integration supports both evaluation via OpenAI-compatible servers and full RL training with exact token IDs through VLLM ManagedServer.

## Core Architecture: From Atropos to Hermes

The integration is built around a three-layer inheritance chain that maps Atropos primitives to Hermes capabilities:

```

atroposlib.BaseEnv  →  HermesAgentBaseEnv  →  ConcreteEnv

```

**`atroposlib.BaseEnv`** provides the generic RL-training server, worker pool, and WandB logging. **`HermesAgentBaseEnv`** ([`environments/hermes_base_env.py`](https://github.com/NousResearch/hermes-agent/blob/main/environments/hermes_base_env.py)) adds the Hermes-specific plumbing: it configures the terminal sandbox backend (local, Docker, Modal), resolves toolsets for each rollout group, and runs the **HermesAgentLoop** inside an Atropos rollout. Concrete environments (e.g., `TerminalTestEnv`, `HermesSweEnv`, or user-defined subclasses) implement five abstract methods required by Atropos: `setup`, `get_next_item`, `format_prompt`, `compute_reward`, and `evaluate`.

The base class exposes a **`ToolContext`** ([`environments/tool_context.py`](https://github.com/NousResearch/hermes-agent/blob/main/environments/tool_context.py)) that gives reward functions direct access to the exact sandbox the model used during the rollout. This ensures deterministic verification—reward functions can invoke terminal, file, or web tools to check the model's work in the same environment that produced it.

## Two-Phase Rollout System

Hermes Agent supports two operating modes depending on your training stage:

**Phase 1 – OpenAI-compatible server** targets evaluation, SFT data generation, and quick benchmarks. The server (OpenAI, OpenRouter, or VLLM) parses tool calls natively; Hermes sends `tools=` and receives structured `tool_calls`. Placeholder token IDs are generated for the Atropos pipeline, making this mode fast but unsuitable for RL algorithms requiring exact log-probs.

**Phase 2 – VLLM ManagedServer** enables full RL training with GRPO or PPO. The server returns raw text; Hermes uses a **client-side tool-call parser** (e.g., `hermes`, `mistral`, `qwen`) to reconstruct `tool_calls` from the model output. Real token IDs flow through Atropos, enabling precise advantage computation. The parser selection is controlled by the `tool_call_parser` configuration field.

## Environment Configuration and Toolset Resolution

### HermesAgentEnvConfig

`HermesAgentBaseEnv` expects a configuration subclass of `HermesAgentEnvConfig` (defined in [`environments/hermes_base_env.py`](https://github.com/NousResearch/hermes-agent/blob/main/environments/hermes_base_env.py)). This config controls:

- **Toolset selection** – `enabled_toolsets`, `disabled_toolsets`, or a probabilistic `distribution` resolved via `toolset_distributions.sample_toolsets_from_distribution`.
- **Terminal backend** – `terminal_backend` (`local`, `docker`, `modal`). The selected backend is exported to environment variables (`TERMINAL_ENV`, `TERMINAL_TIMEOUT`, `TERMINAL_LIFETIME`) so all Hermes tools pick it up automatically.
- **Rollout limits** – `max_agent_turns`, `agent_temperature`, `tool_pool_size`.
- **Phase-2 parser** – `tool_call_parser` (default `hermes`).

### Toolset Resolution

Before each group of rollouts, `_resolve_tools_for_group()` calls `model_tools.get_tool_definitions()` to generate the tool schema:

```python
tools, valid_names = get_tool_definitions(
    enabled_toolsets=group_toolsets,
    disabled_toolsets=self.config.disabled_toolsets,
    quiet_mode=True,
)

```

This resolution happens **once per group**, not per trajectory, reducing overhead significantly. The returned schema list is passed to `HermesAgentLoop` so the LLM knows which tools are available for the current rollout.

## Rollout Execution and Reward Computation

### The collect_trajectory Method

`collect_trajectory()` (lines 5110-5230 in [`hermes_base_env.py`](https://github.com/NousResearch/hermes-agent/blob/main/hermes_base_env.py)) performs the core rollout logic:

1. **Create a unique `task_id`** – guarantees sandbox isolation across parallel workers.
2. **Prepare the initial message list** – combines the optional system prompt with the formatted user prompt from `format_prompt(item)`.
3. **Choose the phase** – `_use_managed_server()` decides whether to use a ManagedServer (Phase 2) or direct server (Phase 1). If Phase 2, it loads the selected parser from `environments/tool_call_parsers.get_parser` and enters a `managed_server` context.
4. **Run the agent loop** – `HermesAgentLoop(...).run(messages)` (from [`environments/agent_loop.py`](https://github.com/NousResearch/hermes-agent/blob/main/environments/agent_loop.py)) handles tool calls, executes them via a `ThreadPoolExecutor(self.config.tool_pool_size)` to avoid deadlocks with async back-ends, and returns an `AgentResult`.
5. **Compute reward** – `compute_reward(item, result, ctx)` receives a fresh `ToolContext(task_id)` pointing at the exact sandbox used during the rollout. Reward functions can invoke any Hermes tool to verify the model's work.
6. **Package the trajectory** – If the server produced real tokens (Phase 2), they are taken from `result.managed_state["nodes"]`; otherwise, placeholder tokens are generated from the conversation text.

### Async-Safety Patches

Many Hermes tools (e.g., the Modal backend) spawn their own event loops via `asyncio.run()`. When called from inside Atropos' event loop, this raises `RuntimeError: Cannot run the event loop while another loop is running`. [`environments/patches.py`](https://github.com/NousResearch/hermes-agent/blob/main/environments/patches.py) (lines 44-86) replaces those calls with a dedicated background thread (`_AsyncWorker`) that runs its own loop, allowing safe execution during rollouts.

### Logging and Visualization

`add_rollouts_for_wandb()` formats the full message trace into a compact, human-readable view via `_format_trajectory_for_display`. This appears in the WandB rollout table and includes tool calls, reasoning steps, and execution results, making debugging RL experiments straightforward.

## Building a Custom Atropos Environment

Below is a minimal example that trains a model to generate Python functions and scores it by running the produced script in the same sandbox.

```python

# my_env.py

from environments.hermes_base_env import HermesAgentBaseEnv, HermesAgentEnvConfig
from atroposlib.envs.server_handling.server_manager import APIServerConfig

class MyEnvConfig(HermesAgentEnvConfig):
    """Custom config – enable only terminal + file tools, use Modal sandbox."""
    pass  # all defaults are fine; you could set terminal_backend="modal" here

class MyEnv(HermesAgentBaseEnv):
    """Simple RL environment that asks the model to write a function."""
    name = "my-env"
    env_config_cls = MyEnvConfig

    @classmethod
    def config_init(cls):
        # Enable the terminal and file toolsets, run in Modal for isolation.

        env_cfg = MyEnvConfig(
            enabled_toolsets=["terminal", "file"],
            terminal_backend="modal",
            max_agent_turns=20,
            tool_pool_size=64,
        )
        # Use an OpenAI‑compatible server (OpenRouter in this example)

        server_cfg = [
            APIServerConfig(
                base_url="https://openrouter.ai/api/v1",
                model_name="anthropic/claude-sonnet-4.6",
                server_type="openai",
            )
        ]
        return env_cfg, server_cfg

    async def setup(self):
        # No external dataset – generate a trivial list of prompts.

        self.prompts = [
            "Write a function `add(a, b)` that returns a + b.",
            "Write a function `reverse(s)` that returns the reversed string.",
        ]
        self.idx = 0

    async def get_next_item(self):
        item = {"prompt": self.prompts[self.idx % len(self.prompts)]}
        self.idx += 1
        return item

    def format_prompt(self, item):
        # The model sees a plain user message.

        return item["prompt"]

    async def compute_reward(self, item, result, ctx):
        # Run the generated script in the same sandbox and check its output.

        # 1. Write the model's answer to a file.

        ctx.write_file("/workspace/solution.py", result.messages[-1]["content"])
        # 2. Execute it with the terminal tool.

        exec_res = ctx.terminal("python /workspace/solution.py")
        # 3. Reward is 1.0 if the script exits cleanly, otherwise 0.0.

        return 1.0 if exec_res["exit_code"] == 0 else 0.0

    async def evaluate(self, *_, **__):
        # Optional periodic evaluation – could run a few held‑out prompts.

        pass

if __name__ == "__main__":
    # CLI entry point – supports `serve`, `process`, `evaluate`

    MyEnv.cli()

```

**Key implementation details**:

- **ToolContext** (`ctx`) gives direct access to the sandbox the model used via `ctx.write_file()` and `ctx.terminal()`.
- The terminal backend is set to **Modal**, which runs each rollout in an isolated cloud container.
- No external dataset is required; the environment supplies its own prompts through `setup()` and `get_next_item()`.

## Running the Integration

Start the Atropos server and connect your environment:

```bash

# 1️⃣  Start the Atropos API (run in a separate terminal)

run-api

# 2️⃣  Run the environment in RL mode (connects to Atropos)

python my_env.py serve \
    --openai.model_name anthropic/claude-opus-4.6 \
    --env.max_agent_turns 30

```

- Use `process` instead of `serve` to generate SFT trajectories without an Atropos server.
- Use `evaluate` to run a benchmark-style evaluation loop.

## Summary

- **HermesAgentBaseEnv** ([`environments/hermes_base_env.py`](https://github.com/NousResearch/hermes-agent/blob/main/environments/hermes_base_env.py)) bridges the Atropos RL server to the Hermes tool-calling loop via a three-layer inheritance chain.
- **ToolContext** ensures reward functions can re-use the exact sandbox that produced the rollout, guaranteeing deterministic verification.
- **Phase 1 vs Phase 2** lets you choose between speed (OpenAI-compatible servers with placeholder tokens) and fidelity (VLLM ManagedServer with real token IDs and log-probs for GRPO/PPO).
- **Async patches** in [`environments/patches.py`](https://github.com/NousResearch/hermes-agent/blob/main/environments/patches.py) enable safe use of Modal and Docker backends inside Atropos' event loop.
- **HermesAgentLoop** ([`environments/agent_loop.py`](https://github.com/NousResearch/hermes-agent/blob/main/environments/agent_loop.py)) provides the reusable multi-turn agent engine that executes tool calls during rollouts.

## Frequently Asked Questions

### How does Hermes Agent handle tool-calling during RL training with Atropos?

During Phase 2 training, Hermes Agent uses client-side parsers in `environments/tool_call_parsers/` to reconstruct tool calls from raw model text generated by the VLLM ManagedServer. This allows Atropos to record exact token IDs and log-probs while Hermes still executes tools in the sandboxed environment. The `HermesAgentLoop` manages the multi-turn conversation, executing tools via a thread pool to prevent deadlocks with async backends.

### What is the difference between Phase 1 and Phase 2 rollouts?

Phase 1 uses standard OpenAI-compatible servers (OpenAI, OpenRouter, or VLLM in chat mode) where the server parses tool calls natively. Hermes receives structured `tool_calls` directly, but Atropos only sees placeholder token IDs, making this suitable for evaluation and SFT data generation. Phase 2 uses VLLM ManagedServer returning raw text; Hermes parses tool calls client-side, allowing Atropos to capture real token IDs and compute precise advantages for RL algorithms like GRPO or PPO.

### How can reward functions access the sandbox environment?

Reward functions receive a `ToolContext` instance (from [`environments/tool_context.py`](https://github.com/NousResearch/hermes-agent/blob/main/environments/tool_context.py)) via the `compute_reward(item, result, ctx)` method. This context provides methods like `ctx.terminal()`, `ctx.write_file()`, and `ctx.read_file()` that point to the exact sandbox container (local, Docker, or Modal) used during the rollout. This design ensures that verification happens in the same environment where the model executed its tools, eliminating inconsistencies.

### Which files should I modify to create a custom RL environment?

Start by subclassing `HermesAgentBaseEnv` in [`environments/hermes_base_env.py`](https://github.com/NousResearch/hermes-agent/blob/main/environments/hermes_base_env.py) and implementing the five required methods: `setup`, `get_next_item`, `format_prompt`, `compute_reward`, and `evaluate`. Define your configuration in a subclass of `HermesAgentEnvConfig` to control toolsets and terminal backends. If you need custom tool-call parsing for Phase 2 training, add a parser to `environments/tool_call_parsers/`. Refer to [`website/docs/developer-guide/environments.md`](https://github.com/NousResearch/hermes-agent/blob/main/website/docs/developer-guide/environments.md) for the full architecture specification.