How to Implement a Custom Environment Server-Side in OpenEnv: reset(), step(), and state() Methods
To implement a custom environment server-side in OpenEnv, subclass the Environment abstract class from openenv.core.env_server.interfaces, define Pydantic-based Action and Observation models, and implement the reset(), step(), and state() methods to handle episode initialization, state transitions, and current state inspection.
OpenEnv, developed by Hugging Face, provides a standardized framework for serving interactive reinforcement learning environments via HTTP or WebSocket. To implement a custom environment server-side with reset(), step(), and state() methods in OpenEnv, you must create a concrete class that inherits from the abstract Environment base class and define the data contracts that govern client-server communication.
Understanding the Core Environment Interface
The foundation of any server-side environment is the abstract Environment class defined in src/openenv/core/env_server/interfaces.py at line 55. This contract mandates three core operations: episode initialization via reset(), state progression via step(), and state inspection via the state property.
Required Method Signatures
When subclassing Environment, you must implement these specific signatures exactly as defined in the source:
reset(self, seed: Optional[int] = None, episode_id: Optional[str] = None, **kwargs) -> Observation: Reinitializes the environment state and returns an initial observation.step(self, action: Action, timeout_s: Optional[float] = None, **kwargs) -> Observation: Processes an action, updates internal state, computes rewards, and returns the resulting observation.@property state(self) -> State: Returns the currentStateobject containing episode metadata such as step count and episode ID.
Step-by-Step Implementation Guide
Define Action and Observation Models
First, create Pydantic models that inherit from openenv.core.env_server.types.Action and Observation. These define the JSON schema for client requests and server responses according to the types defined in src/openenv/core/env_server/types.py.
# src/my_env/models.py
from openenv.core.env_server.types import Action, Observation
from pydantic import Field
class EchoAction(Action):
"""Message sent by the client."""
message: str = Field(..., description="Message to echo back")
class EchoObservation(Observation):
"""Result returned after each step."""
echoed_message: str = Field(..., description="The echoed message")
message_length: int = Field(..., description="Length of the message")
Create the Concrete Environment Class
Implement a concrete class that subclasses Environment and provides the three required methods. Store internal state in a State object typically named self._state and expose it via the state property.
# src/my_env/server/echo_environment.py
from uuid import uuid4
from openenv.core.env_server.interfaces import Environment
from openenv.core.env_server.types import State
from ..models import EchoAction, EchoObservation
class EchoEnvironment(Environment):
"""A simple echo environment demonstrating the required interface."""
# Enable per-client isolation (optional)
SUPPORTS_CONCURRENT_SESSIONS: bool = True
def __init__(self):
self._state = State(episode_id=str(uuid4()), step_count=0)
self._reset_count = 0
def reset(self, seed: int | None = None, episode_id: str | None = None, **kwargs) -> EchoObservation:
"""Reset the environment to initial state."""
# Re-initialize internal state
self._state = State(episode_id=str(uuid4()), step_count=0)
self._reset_count += 1
# Clear rubric trajectory if using automated reward computation
self._reset_rubric()
return EchoObservation(
echoed_message="Echo environment ready!",
message_length=0,
done=False,
reward=0.0,
)
def step(self, action: EchoAction, timeout_s: float | None = None, **kwargs) -> EchoObservation:
"""Process one step of the environment."""
self._state.step_count += 1
length = len(action.message)
# Simple reward computation based on message length
reward = length * 0.1
return EchoObservation(
echoed_message=action.message,
message_length=length,
done=False,
reward=reward,
metadata={"step": self._state.step_count},
)
@property
def state(self) -> State:
"""Return current environment state."""
return self._state
Wire the Environment into the Server
Expose the environment via HTTP using the make_http_app factory from src/openenv/core/env_server/http_server.py. This creates FastAPI endpoints that serialize your Observation returns into ResetResponse or StepResponse models.
# src/my_env/server/app.py
from fastapi import FastAPI
from openenv.core.env_server.http_server import make_http_app
from .echo_environment import EchoEnvironment
app: FastAPI = make_http_app(
environment_factory=EchoEnvironment,
# optional: provide a Rubric instance here for automated scoring
)
Key Implementation Details
State Management and Concurrency
Track episode state using a private State attribute (self._state) and expose it through the state property. Set SUPPORTS_CONCURRENT_SESSIONS = True if your environment isolates state per client, allowing multiple simultaneous episodes. When False (the default), the server assumes single-session semantics according to the implementation in src/openenv/core/env_server/interfaces.py.
Rubric Integration for Reward Computation
OpenEnv supports automated reward computation through Rubrics. Call self._reset_rubric() at the start of reset() to clear previous trajectories, and optionally use self._apply_rubric(action, observation) inside step() to delegate reward calculation to a configured rubric instance rather than computing it manually.
Registration via CLI Templates
Register new environments using the CLI template system so that openenv serve can discover them. The template files live under src/openenv/cli/templates/openenv_env/. Running openenv init copies these templates into a new folder, replacing placeholders like __ENV_CLASS_NAME__ with your specific implementation names.
Serialization and HTTP Endpoints
The framework automatically serializes your Observation subclasses to JSON via Pydantic. The HTTP server defined in src/openenv/core/env_server/http_server.py exposes:
POST /reset→ invokesreset()and returns aResetResponsePOST /step→ invokesstep()and returns aStepResponseGET /state→ returns the currentStateobject
WebSocket support is handled by src/openenv/core/env_server/web_interface.py, which manages message-based communication for the same three operations.
Summary
- Subclass
Environment: Inherit fromopenenv.core.env_server.interfaces.Environmentto establish the required contract forreset(),step(), andstate. - Implement the three core methods: Define episode initialization, action processing, and state inspection to handle the full environment lifecycle.
- Define data models: Create Pydantic subclasses of
ActionandObservationfromopenenv.core.env_server.typesfor type-safe client-server communication. - Configure concurrency: Set
SUPPORTS_CONCURRENT_SESSIONSbased on whether your environment supports isolated simultaneous sessions. - Use Rubric methods: Call
_reset_rubric()and_apply_rubric()if leveraging automated reward computation. - Initialize with CLI: Use the templates in
src/openenv/cli/templates/openenv_env/viaopenenv initto scaffold new environments.
Frequently Asked Questions
What is the difference between the state property and the State type in OpenEnv?
The state property is a method you implement in your concrete environment class that returns a State object. The State type, defined in src/openenv/core/env_server/types.py, is a Pydantic model that tracks episode metadata such as episode_id and step_count. This distinction separates the interface used to expose state from the internal data container itself.
How do I handle concurrent sessions in my custom OpenEnv environment?
Set the class attribute SUPPORTS_CONCURRENT_SESSIONS = True if your environment can isolate state between multiple simultaneous clients. When enabled, the server infrastructure in src/openenv/core/env_server/http_server.py manages separate state instances per session. If set to False (the default), the server assumes single-session semantics and will not isolate client interactions.
Can I customize the reward computation logic in OpenEnv environments?
Yes. You can compute rewards manually within the step() method by setting the reward field on your returned Observation subclass, or you can integrate with OpenEnv's Rubric system. Call self._apply_rubric(action, observation) to delegate scoring to a configurable rubric instance, and call self._reset_rubric() during reset() to clear accumulated trajectory data before starting a new episode.
Where are the HTTP endpoints for reset and step defined in the OpenEnv source code?
The HTTP endpoints are defined in src/openenv/core/env_server/http_server.py, which maps POST /reset to your environment's reset() method and POST /step to your step() method. WebSocket equivalents are implemented in src/openenv/core/env_server/web_interface.py. Both serialize your Observation returns into ResetResponse or StepResponse models defined in src/openenv/core/env_server/types.py.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →