How to Write Unit Tests for Custom OpenEnv Environment Implementations

OpenEnv provides typed EnvClient classes and synchronous wrappers that enable you to write pytest-based unit tests for custom environments by launching Docker images, stepping through actions, and asserting on StepResult observations, rewards, and termination flags.

OpenEnv is a Hugging Face framework for developing and deploying reproducible environment implementations. When you build a custom environment, you must validate that your action/observation protocol, state initialization, and reward logic function correctly across interaction cycles. This guide demonstrates how to write unit tests for custom OpenEnv environment implementations using the core client abstractions and testing patterns found in the official repository.

Core Testing Abstractions in OpenEnv

The OpenEnv architecture provides three primary client classes for testing:

  • EnvClient[Action, Observation, State]: A generic typed client that you subclass for your specific environment types.
  • GenericEnvClient: A loosely-typed client defined in src/openenv/core/generic_client.py that works with any OpenEnv-compliant environment.
  • SyncEnvClient: A synchronous wrapper around async clients located in src/openenv/core/sync_client.py that simplifies test syntax.

When implementing tests, you typically instantiate your typed client subclass (e.g., MyEnv(EnvClient[MyAction, MyObservation, MyState])) and interact with it through the standardized lifecycle methods.

The Four-Phase Testing Pattern

According to the test suite in tests/test_core/test_generic_client.py, effective OpenEnv unit tests follow a consistent four-phase pattern:

  1. Instantiate: Launch the environment using from_docker_image() or from_env().
  2. Reset: Call client.reset() to verify initial state configuration.
  3. Step: Execute client.step(action) and inspect the returned StepResult.
  4. Assert: Validate observation data, reward scalars, and termination flags.

This pattern ensures that your Docker image builds correctly, your protocol implementation responds to actions, and your reward rubric returns expected values for specific trajectories.

Writing Async Unit Tests

For environments requiring asynchronous interaction, use the native async client methods. The following example demonstrates the recommended pattern for writing unit tests for custom OpenEnv environment implementations:


# tests/envs/test_my_env.py

import pytest
from my_env.client import MyEnv
from my_env.models import MyAction

@pytest.mark.asyncio
async def test_my_env_basic_flow():
    # Launch environment from Docker image

    async with MyEnv.from_docker_image("my-env:latest") as client:
        # Verify initial state

        init_state = await client.reset()
        assert init_state.some_flag is True
        
        # Execute action and inspect result

        action = MyAction(command="do_something", parameters={"x": 1})
        step_result = await client.step(action)
        
        # Validate StepResult components

        assert step_result.observation.result == "expected output"
        assert step_result.reward == pytest.approx(1.0)
        assert step_result.done is False
        
        # Test terminal conditions

        while not step_result.done:
            step_result = await client.step(action)
        assert step_result.reward == pytest.approx(10.0)

The async with context manager ensures proper resource cleanup, while assertions on step_result validate the observation content, reward calculation, and episode termination logic implemented in your environment's rubric.

Writing Synchronous Unit Tests with SyncEnvClient

When you prefer traditional synchronous test syntax, use the sync() method provided by src/openenv/core/sync_client.py. This eliminates the need for async/await keywords while maintaining identical functionality:


# Synchronous alternative using SyncEnvClient

def test_my_env_sync():
    with MyEnv.from_docker_image("my-env:latest").sync() as client:
        # Reset and verify state

        init_state = client.reset()
        assert init_state.some_flag is True
        
        # Step through environment

        action = MyAction(command="do_something", parameters={"x": 1})
        step = client.step(action)
        
        # Assert on StepResult

        assert step.observation.result == "expected output"
        assert step.reward == pytest.approx(1.0)
        assert not step.done

SyncEnvClient wraps the async implementation, making it ideal for straightforward test suites that don't require concurrent environment management.

Reference Files and Implementation Details

The following files in the Hugging Face OpenEnv repository contain the reference implementations and patterns for unit testing:

Summary

  • OpenEnv environments are tested through typed EnvClient subclasses that handle the underlying protocol.
  • Use from_docker_image() in tests to validate that your environment builds and runs correctly in its containerized form.
  • Always test the full lifecycle: reset() for initialization, step() for transitions, and assertions on StepResult for observations, rewards, and done flags.
  • Choose async or sync based on your test suite needs: native async for performance, SyncEnvClient for simplicity.
  • Reference tests/test_core/test_generic_client.py for canonical patterns on exercising custom environment implementations.

Frequently Asked Questions

How do I test my OpenEnv environment without building a Docker image every time?

Use MyEnv.from_env() instead of from_docker_image() to connect to an already-running local instance or in-process environment. This speeds up iteration during development while still allowing you to test the full client protocol against your custom implementation.

What should I assert about the StepResult in my unit tests?

Assert on three specific fields: step_result.observation to verify your environment returns correct state data, step_result.reward using pytest.approx() for floating-point comparisons, and step_result.done to confirm episode termination logic works correctly for your environment's success conditions.

Can I use standard pytest fixtures with OpenEnv clients?

Yes. Create fixtures that yield clients using with MyEnv.from_docker_image(...).sync() as client for synchronous tests or async with MyEnv.from_docker_image(...) as client for async tests. Ensure proper cleanup by using context managers rather than bare instantiations to prevent resource leaks between test cases.

How do I test custom reward logic in my OpenEnv implementation?

Execute specific action sequences that should trigger known reward values, then assert that step_result.reward matches expected scalars. For complex rubrics, test edge cases like terminal states (where done=True) separately from intermediate steps, as documented in docs/source/guides/rewards.md.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →