Implementing Multi-Agent Environments with Shared or Separate State Using OpenEnv

OpenEnv provides a Gymnasium-style abstraction where you implement multi-agent environments by subclassing Environment and managing either a single shared State object or per-agent private sub-states, with the EnvClient handling async communication over WebSocket or HTTP.

OpenEnv from Hugging Face offers a clean, production-ready framework for building agentic execution environments. Whether you are simulating competitive games with full observability or training independent agents with hidden information, the library provides distinct patterns for implementing multi-agent environments with shared or separate state. This guide examines the core architecture, concrete implementation patterns, and key source files in the huggingface/OpenEnv repository.

Core Architecture Components

OpenEnv separates concerns between the simulation server, typed data models, and the client interface. Understanding these components is essential before implementing either state pattern.

The Environment Server

In src/openenv/core/env_server/mcp_environment.py, the abstract Environment class defines the contract your multi-agent simulation must implement. Subclasses override reset() to initialize the global state and step(action) to process agent moves. The server maintains the authoritative state and exposes it via WebSocket or HTTP endpoints defined in src/openenv/core/env_server/http_server.py.

Data Models

Typed dataclasses in src/openenv/core/models.py (or your environment-specific models.py) define the contract between client and server. You must define:

  • Action: Contains the agent's move and an agent_id field
  • Observation: What the agent perceives after a step
  • State: The complete internal representation of the world

Refer to envs/echo_env/models.py for minimal typing conventions.

The Client Interface

The EnvClient in src/openenv/core/env_client.py provides the thin async wrapper that serializes actions and deserializes observations. Generated environment clients inherit from this class, automatically handling the wire protocol so your agent code remains agnostic to whether the server runs locally or in a Docker container managed by src/openenv/core/containers/runtime/providers.py.

Shared-State Multi-Agent Pattern

In the shared-state pattern, all agents read from and write to a single global State object. This approach suits fully observable games like chess or collaborative grid worlds.

Defining the Global State

Create a dataclass that holds all information visible to every agent. In your environment's models.py, define a structure that tracks per-agent positions within a global board:

from dataclasses import dataclass
from typing import Dict, Tuple

@dataclass
class MultiAgentState:
    board: list[list[int]]
    positions: Dict[str, Tuple[int, int]]
    step_count: int = 0

Server Implementation

Subclass Environment in your server file to manage the shared state. The step method receives one action at a time and updates the global structures:

from openenv.core.env_server import Environment
from .models import GridAction, GridObservation, MultiAgentState

class GridEnvironment(Environment):
    def __init__(self):
        super().__init__()
        self.state = MultiAgentState(
            board=[[0]*5 for _ in range(5)],
            positions={"A": (0, 0), "B": (4, 4)}
        )

    async def reset(self) -> GridObservation:
        self.state.step_count = 0
        return GridObservation(
            board=self.state.board,
            positions=self.state.positions
        )

    async def step(self, action: GridAction) -> GridObservation:
        x, y = self.state.positions[action.agent_id]
        nx, ny = x + action.dx, y + action.dy
        if 0 <= nx < 5 and 0 <= ny < 5:
            self.state.positions[action.agent_id] = (nx, ny)
        self.state.step_count += 1
        return GridObservation(
            board=self.state.board,
            positions=self.state.positions
        )

Client Usage

Instantiate a single EnvClient (or your generated wrapper) and alternate moves by changing the agent_id in the Action:

import asyncio
from my_grid_env.client import GridEnv, GridAction

async def play():
    async with GridEnv(base_url="http://localhost:8000") as env:
        await env.reset()
        await env.step(GridAction(agent_id="A", dx=1, dy=0))
        await env.step(GridAction(agent_id="B", dx=-1, dy=0))
        obs = await env.step(GridAction(agent_id="A", dx=0, dy=1))
        print("Shared board:", obs.board)

asyncio.run(play())

Because the server holds one authoritative state, both agents automatically see each other's updates.

Separate-State Multi-Agent Pattern

When agents require private information—such as hidden cards in poker or independent episodes—use the separate-state pattern. Here, the server maintains a root state containing private sub-states for each agent.

Structuring Private and Public Data

Define a root state that splits global information from agent-specific slices:

from dataclasses import dataclass
from typing import Dict, Tuple

@dataclass
class PrivateAgentState:
    position: Tuple[int, int]
    health: int

@dataclass
class MultiAgentRootState:
    global_score: int
    private: Dict[str, PrivateAgentState]

Isolating Agent Views

The server updates only the requesting agent's private slice during step, as implemented in src/openenv/core/env_server/mcp_environment.py:

class PrivateGridEnvironment(Environment):
    def __init__(self):
        super().__init__()
        self.state = MultiAgentRootState(
            global_score=0,
            private={
                "A": PrivateAgentState((0, 0), 10),
                "B": PrivateAgentState((4, 4), 10)
            }
        )

    async def step(self, action: GridAction) -> GridObservation:
        p_state = self.state.private[action.agent_id]
        nx, ny = p_state.position[0] + action.dx, p_state.position[1] + action.dy
        if 0 <= nx < 5 and 0 <= ny < 5:
            p_state.position = (nx, ny)
        self.state.global_score += 1
        return GridObservation(
            global_score=self.state.global_score,
            private_state=p_state
        )

Parallel Agent Execution

Clients run in separate processes or threads, each identifying itself via the agent_id parameter. The server verifies this ID (potentially via MCP authentication) and returns only the matching private slice:

import asyncio
from my_private_env.client import PrivateGridEnv, GridAction

async def play_one_agent(agent_id: str):
    async with PrivateGridEnv(
        base_url="http://localhost:8000",
        agent_id=agent_id
    ) as env:
        await env.reset()
        await env.step(GridAction(dx=1, dy=0))
        obs = await env.step(GridAction(dx=0, dy=1))
        print(f"Agent {agent_id} position:", obs.private_state.position)

asyncio.run(asyncio.gather(
    play_one_agent("A"),
    play_one_agent("B")
))

This pattern ensures agents cannot observe each other's private states, making it ideal for imperfect information games.

Deployment and Tool Integration

OpenEnv provides CLI utilities to package your environment for local or distributed execution.

Scaffolding and Containerization

Generate a new environment skeleton and build the container:

openenv init my_multi_agent_env
openenv build
openenv serve

The openenv serve command launches a FastAPI server locally for rapid iteration. For production, container providers in src/openenv/core/containers/runtime/providers.py abstract Docker details for Swarm or Kubernetes deployment.

Enabling MCP for LLM Agents

When agents are large language models that invoke tools, enable the Model-Context-Protocol:

export ENABLE_MCP=true

With MCP enabled, the environment acts as a tool provider, streaming function calls defined in src/openenv/core/env_server/mcp_types.py. The client forwards LLM-generated tool calls to the environment, receives results, and returns them to the model, enabling production-grade agentic workflows.

Summary

  • OpenEnv implements a client-server architecture where Environment subclasses in src/openenv/core/env_server/mcp_environment.py maintain authoritative state.
  • Shared-state environments store a single State object accessible to all agents, suitable for fully observable simulations.
  • Separate-state environments partition private data per agent using sub-state dictionaries, necessary for hidden information games.
  • The EnvClient in src/openenv/core/env_client.py abstracts WebSocket communication, allowing synchronous or asynchronous agent control.
  • Use openenv init, build, and serve to scaffold and deploy containers, and enable MCP support when integrating LLM tool calling.

Frequently Asked Questions

What is the difference between shared-state and separate-state multi-agent environments?

Shared-state environments maintain one global State object that all agents observe and modify, making every agent's view identical to the server's view. Separate-state environments store private sub-states for each agent ID and return only the relevant slice to each caller, enabling hidden information and independent episode tracking.

How does communication work between the client and the OpenEnv server?

The EnvClient in src/openenv/core/env_client.py establishes a WebSocket or HTTP connection to the server. It serializes Action dataclasses to JSON, transmits them via await client.step(action), and deserializes the returned Observation. This async pattern allows multiple agents to connect simultaneously from different processes.

Can I use synchronous clients with OpenEnv multi-agent environments?

Yes. While the underlying EnvClient uses async/await, you can wrap it with the .sync() method to obtain a synchronous interface. This is useful for blocking agent policies or integration with synchronous training loops, as shown in the separate-state client examples.

When should I enable MCP support in OpenEnv?

Enable MCP by setting ENABLE_MCP=true when your agents are large language models that need to invoke tools (such as code execution or search functions). The Model-Context-Protocol allows the environment to act as a tool provider, streaming structured tool calls between the LLM and the simulation server according to the types defined in src/openenv/core/env_server/mcp_types.py.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →