# OpenEnv Development Roadmap: Three Phases to Production-Grade Evaluation

> Explore the OpenEnv development roadmap, detailing three phases: environment abstraction, reward pipelines, and evaluation support. See how OpenEnv evolves into a production-grade agent evaluation platform.

- Repository: [Hugging Face/OpenEnv](https://github.com/huggingface/OpenEnv)
- Tags: architecture
- Published: 2026-06-16

---

**OpenEnv is advancing through three incremental phases—environment abstraction (0.3), reward pipelines (0.6), and evaluation support (0.9→1.0)—to evolve from a minimal sandboxed API into a full-featured agent evaluation platform.**

The `huggingface/OpenEnv` repository provides a sandboxed environment API designed for training and evaluating AI agents. According to the canonical RFC document [[`rfcs/000-project-phases.md`](https://github.com/huggingface/OpenEnv/blob/main/rfcs/000-project-phases.md)](https://github.com/huggingface/OpenEnv/blob/main/rfcs/000-project-phases.md), the development roadmap follows a deliberate trajectory that prioritizes stable foundations before introducing complex reward mechanisms and evaluation frameworks.

## Phase 1: Environment Abstraction and Tool Support (v0.3)

Phase 1 targets version **0.3** as the first public release, focusing on defining the core **environment abstraction** and establishing robust sandboxing capabilities. This phase establishes the conventions for what constitutes an environment, dependency distribution plumbing, and universal tool interfaces.

The key milestones in this phase include:

- **Environment conventions** (RFC 001) that define state management and episode lifecycle
- **Dependency and distribution** infrastructure (RFC 002) for binary packaging
- **Traditional tool calls** via `ListToolsAction` and `CallToolAction` (RFC 003)
- **CodeAct support** enabling agents to write and execute arbitrary code (RFC 004)
- **Model Context Protocol (MCP)** integration (RFC 005) as the universal interface for tool discovery

According to the source code in [[`src/openenv_core/__init__.py`](https://github.com/huggingface/OpenEnv/blob/main/src/openenv_core/__init__.py)](https://github.com/huggingface/OpenEnv/blob/main/src/openenv_core/__init__.py) and [[`src/openenv/__init__.py`](https://github.com/huggingface/OpenEnv/blob/main/src/openenv/__init__.py)](https://github.com/huggingface/OpenEnv/blob/main/src/openenv/__init__.py), this phase solidifies the **client-server separation** pattern where the thin `EnvClient` connects over WebSockets to a Dockerized FastAPI server.

## Phase 2: Reward Pipelines and External Services (v0.6)

Phase 2 targets version **0.6** and introduces **reward pipelines** alongside the first RPC mechanisms. This phase builds upon the stable foundation from Phase 1, allowing environments to call out to external reward models for learned feedback signals.

The implementation focuses on designing version-controlled reward interfaces and RPC extensions for external services. The architectural foundation for this phase is documented in [[`rfcs/004-rubrics.md`](https://github.com/huggingface/OpenEnv/blob/main/rfcs/004-rubrics.md)](https://github.com/huggingface/OpenEnv/blob/main/rfcs/004-rubrics.md), which outlines reward-pipeline design considerations.

The **modular container provider** architecture—implemented in [[`src/openenv_core/providers.py`](https://github.com/huggingface/OpenEnv/blob/main/src/openenv_core/providers.py)](https://github.com/huggingface/OpenEnv/blob/main/src/openenv_core/providers.py)—supports `LocalDockerProvider`, `DockerSwarmProvider`, and `KubernetesProvider` configurations that will host reward-specific containers in this phase.

## Phase 3: Evaluation Support and Stable Release (v0.9→1.0)

Phase 3 targets versions **0.9 through 1.0**, adding comprehensive **evaluation support** and finalizing the platform for production-grade usage. This phase introduces three distinct evaluation modalities: simple scripted evals, LLM-judge evals, and fully-agentic evaluation orchestration.

The milestones include implementing script-based evaluation runners, LLM-as-judge scoring mechanisms, and end-to-end agentic evaluation orchestration. The phase concludes with stability guarantees and API finalization for the 1.0 release.

## Architectural Foundations Supporting the Roadmap

### Client-Server Separation

The architecture locks down a thin client (`EnvClient`) interacting over WebSockets with a Dockerized FastAPI server. This pattern, visible in [[`src/openenv_core/__init__.py`](https://github.com/huggingface/OpenEnv/blob/main/src/openenv_core/__init__.py)](https://github.com/huggingface/OpenEnv/blob/main/src/openenv_core/__init__.py), ensures that later reward and eval services can be added without breaking existing client implementations.

### Tool Registry and MCP Integration

Phase 1's tool-registry design adds a typed, discoverable set of actions that agents can invoke. The **Model Context Protocol (MCP)** layer abstracts these calls through `CallToolAction` and `ListToolsAction`, as specified in [[`rfcs/003-mcp-support.md`](https://github.com/huggingface/OpenEnv/blob/main/rfcs/003-mcp-support.md)](https://github.com/huggingface/OpenEnv/blob/main/rfcs/003-mcp-support.md), ensuring that reward and evaluation services integrate seamlessly in subsequent phases.

### Modular Container Providers

The provider architecture is deliberately lightweight in Phase 1, supporting `LocalDockerProvider` and distributed variants. This modularity allows Phase 2 to add reward-specific containers and Phase 3 to deploy distributed evaluation workers without architectural rewrites.

### Versioning Discipline

Each phase concludes with a semantic version bump (0.3 → 0.6 → 0.9 → 1.0). The CI pipelines in **`.github/workflows/*.yml`** enforce these version gates, ensuring breaking changes only occur at designated phase boundaries.

## Code Examples

### Current Phase 1 Environment Interaction

```python
import asyncio
from echo_env import CallToolAction, EchoEnv

async def run():
    # Connect to the Echo environment (WebSocket client)

    async with EchoEnv(base_url="https://openenv-echo-env.hf.space") as client:
        # Initialise a new episode

        obs = await client.reset()
        print(obs.echoed_message)          # → "Echo environment ready!"

        # Send a simple tool call

        result = await client.step(
            CallToolAction(tool_name="echo_message", arguments={"message": "Hello, OpenEnv!"})
        )
        print(result.observation.result)   # → "Hello, OpenEnv!"

        print("Reward:", result.reward)    # Reward is currently a placeholder (Phase 1)

asyncio.run(run())

```

This example demonstrates the core `reset` / `step` cycle using the MCP tool interface available in the current Phase 1 release.

### CLI Environment Scaffolding

```bash

# Initialise a fresh environment skeleton

openenv init my_env

```

The CLI creates a directory structure ready for Phase 1 development. Future releases will add flags for reward and evaluation scaffolding as Phases 2 and 3 become available.

### Phase 2 Reward Integration (Preview)

```python

# Pseudo-code for reward-aware step (to be available after Phase 2)

result = await client.step(
    CallToolAction(tool_name="write_code", arguments={"code": "..."}),
    reward_config={"type": "llm_score", "model": "gpt-4o"}
)
print("Reward:", result.reward)  # Returns a learned scalar from the reward pipeline

```

When Phase 2 lands, the `step` method will accept an optional `reward_config` parameter that routes observations through the new reward service.

### Phase 3 Evaluation Loop (Preview)

```python

# High-level evaluation driver (Phase 3)

from openenv.eval import EvaluationRunner

runner = EvaluationRunner(
    env_name="my_env",
    eval_suite="basic_score",
    num_episodes=100,
)

metrics = runner.run()
print(metrics)  # {'average_score': 0.84, 'success_rate': 0.71}

```

The `EvaluationRunner` abstraction will encapsulate scripted, LLM-judge, and agentic evaluation pipelines introduced in Phase 3.

## Summary

- OpenEnv follows a **three-phase roadmap** (0.3 → 0.6 → 1.0) that incrementalizes complexity from basic environment abstraction to full evaluation support.
- **Phase 1** establishes the sandboxed environment API, MCP tool support, and containerized architecture through RFCs 001-005.
- **Phase 2** introduces reward pipelines and RPC mechanisms for external reward models, building on the stable Phase 1 foundation.
- **Phase 3** completes the platform with evaluation runners, LLM-judge capabilities, and agentic orchestration before the stable 1.0 release.
- The architecture in [[`src/openenv_core/__init__.py`](https://github.com/huggingface/OpenEnv/blob/main/src/openenv_core/__init__.py)](https://github.com/huggingface/OpenEnv/blob/main/src/openenv_core/__init__.py) and [[`src/openenv_core/providers.py`](https://github.com/huggingface/OpenEnv/blob/main/src/openenv_core/providers.py)](https://github.com/huggingface/OpenEnv/blob/main/src/openenv_core/providers.py) supports this evolution through client-server separation and modular providers.

## Frequently Asked Questions

### What is the current release phase of OpenEnv?

OpenEnv is currently in **Phase 1**, targeting version 0.3 as its first public release. This phase focuses on environment abstraction, sandboxing, and Model Context Protocol (MCP) tool support as documented in [[`rfcs/000-project-phases.md`](https://github.com/huggingface/OpenEnv/blob/main/rfcs/000-project-phases.md)](https://github.com/huggingface/OpenEnv/blob/main/rfcs/000-project-phases.md).

### How does OpenEnv handle reward calculation in future phases?

**Phase 2** (v0.6) will introduce reward pipelines that allow environments to call external reward models via RPC mechanisms. The `step` method will accept a `reward_config` parameter to route observations through these services, returning learned scalar rewards rather than placeholder values.

### What is the Model Context Protocol (MCP) in OpenEnv?

The **Model Context Protocol (MCP)** is the universal interface layer implemented in Phase 1 that standardizes tool discovery and invocation through typed actions like `CallToolAction` and `ListToolsAction`. Defined in [[`rfcs/003-mcp-support.md`](https://github.com/huggingface/OpenEnv/blob/main/rfcs/003-mcp-support.md)](https://github.com/huggingface/OpenEnv/blob/main/rfcs/003-mcp-support.md), MCP ensures that tools, rewards, and evaluations use a consistent communication protocol.

### When will OpenEnv reach stable 1.0?

OpenEnv targets version **1.0** at the conclusion of Phase 3, following the 0.9 release. This stable release will include comprehensive evaluation support, LLM-judge capabilities, and full agentic evaluation orchestration, with stability guarantees established after the Phase 3 milestones are complete.