# Using OpenEnv with torchforge for Agentic RL Training Workflows

> Leverage OpenEnv with torchforge for agentic RL training. Seamlessly integrate containerized workloads for type-safe reinforcement learning workflows direct from Gymnasium interfaces.

- Repository: [Hugging Face/OpenEnv](https://github.com/huggingface/OpenEnv)
- Tags: tutorial
- Published: 2026-06-14

---

**OpenEnv provides a standardized client-server stack that exposes isolated execution environments as Gymnasium-style interfaces, enabling torchforge's GRPO trainer to call `reset()` and `step()` directly on containerized workloads for type-safe agentic reinforcement learning.**

OpenEnv is a Hugging Face framework that containerizes arbitrary execution environments behind a WebSocket API, making them compatible with standard RL training loops. When paired with torchforge—Meta's PyTorch-based agentic RL library—this architecture enables clean, reproducible GRPO training on complex, tool-driven tasks. This guide explains how to wire OpenEnv's client-server stack into torchforge's training workflows using the concrete Blackjack example from the official repository.

## Architecture Overview

### EnvClient (Client-Side Interface)

In [`src/openenv/core/env_client.py`](https://github.com/huggingface/OpenEnv/blob/main/src/openenv/core/env_client.py) (lines 65-68), the `EnvClient` class handles async WebSocket communication and container lifecycle management. It exposes a synchronous `.sync()` wrapper for blocking trainers, but operates async by default to prevent I/O bottlenecks during high-frequency `step()` calls. The client automatically launches containers via the configured provider and parses Pydantic-typed responses into native Python objects.

### Environment Server

The server-side logic lives in [`src/openenv/core/env_server/mcp_environment.py`](https://github.com/huggingface/OpenEnv/blob/main/src/openenv/core/env_server/mcp_environment.py) (lines 12-18). Concrete environments inherit from the base `Environment` class and implement `reset()`, `step()`, and `state()` methods. These return typed observations and rewards via Pydantic models, ensuring schema validation between the containerized server and the training loop.

### Container Providers

Container orchestration is abstracted in [`src/openenv/core/containers/runtime/providers.py`](https://github.com/huggingface/OpenEnv/blob/main/src/openenv/core/containers/runtime/providers.py) (lines 10-16). Providers handle Docker, Kubernetes, Daytona, or cloud runtimes, spinning up isolated sandboxes that execute the environment server. This lets torchforge trainers treat infrastructure as fungible compute without code changes.

### torchforge GRPO Integration

The torchforge `GRPOTrainer` accepts an `environment_factory` callable that returns OpenEnv clients. According to the repository README (line 337) and the GRPO Blackjack example, the trainer invokes `env.reset()` and `env.step(action)` directly, receiving structured `StepResult` objects containing observations, rewards, and done flags.

## Training Data Flow

The integration follows a five-step pipeline that bridges torchforge's policy updates with OpenEnv's containerized execution:

1. **Trainer Initialization**: The torchforge trainer loads a configuration specifying an `environment_factory` that instantiates OpenEnv clients.
2. **Container Launch**: When the factory creates a client (e.g., `BlackjackEnv`), the container provider spins up a Docker sandbox running the environment server.
3. **Action Execution**: During rollout, the trainer calls `env.step(action)`; the client serializes the action via WebSocket to the container.
4. **Typed Response**: The server executes the step, computes rewards, and returns Pydantic-typed results that torchforge converts into trajectory batches.
5. **Policy Updates**: GRPO computes policy and value updates from the collected trajectories. The optional MCP (Model Context Protocol) layer is bypassed when running in-process, as documented in [`docs/source/tutorials/mcp-environment.md`](https://github.com/huggingface/OpenEnv/blob/main/docs/source/tutorials/mcp-environment.md) (line 46).

## End-to-End Implementation Example

The following script reproduces the Blackjack GRPO workflow from `examples/grpo_blackjack/`. It assumes OpenEnv is built (`openenv build`) and torchforge is installed (`pip install git+https://github.com/meta-pytorch/torchforge.git`).

```python

# train_blackjack_grpo.py

import asyncio
from pathlib import Path

from forge.trainer import GRPOTrainer
from forge.config import LauncherConfig, ProvisionerConfig

from grpo_blackjack.client import BlackjackEnv

def make_env():
    """Factory function returning the OpenEnv client."""
    return BlackjackEnv()

trainer = GRPOTrainer(
    environment_factory=make_env,
    config_path=Path("examples/grpo_blackjack/blackjack.yaml"),
    provisioner=ProvisionerConfig(provider="docker"),
)

async def main():
    await trainer.train(num_iterations=200)

if __name__ == "__main__":
    asyncio.run(main())

```

### Code Breakdown

- **Lines 6-7**: Import torchforge's `GRPOTrainer` and configuration classes as shown in the example README (lines 33-34).
- **Line 9**: Import the generated OpenEnv client from the scaffolded example at [`examples/grpo_blackjack/client.py`](https://github.com/huggingface/OpenEnv/blob/main/examples/grpo_blackjack/client.py).
- **Lines 12-13**: Define the factory that torchforge calls to create fresh environment instances, mirroring the pattern in [`docs/source/tutorials/mcp-environment.md`](https://github.com/huggingface/OpenEnv/blob/main/docs/source/tutorials/mcp-environment.md) (line 46).
- **Lines 15-19**: Configure the trainer with the Docker provider and hyperparameters from [`examples/grpo_blackjack/blackjack.yaml`](https://github.com/huggingface/OpenEnv/blob/main/examples/grpo_blackjack/blackjack.yaml).
- **Lines 23-24**: Launch 200 GRPO iterations, covering approximately 2 million environment steps via the async entry point.

## Key Source Files and References

- **[`src/openenv/core/env_client.py`](https://github.com/huggingface/OpenEnv/blob/main/src/openenv/core/env_client.py)** (lines 65-68): Core WebSocket client implementation with sync wrapper.
- **[`src/openenv/core/env_server/mcp_environment.py`](https://github.com/huggingface/OpenEnv/blob/main/src/openenv/core/env_server/mcp_environment.py)** (lines 12-18): Base environment class providing `reset`/`step` RPC entry points.
- **[`src/openenv/core/containers/runtime/providers.py`](https://github.com/huggingface/OpenEnv/blob/main/src/openenv/core/containers/runtime/providers.py)** (lines 10-16): Abstraction layer for Docker and Kubernetes providers.
- **[`examples/grpo_blackjack/grpo_utils.py`](https://github.com/huggingface/OpenEnv/blob/main/examples/grpo_blackjack/grpo_utils.py)**: Utility functions that convert torchforge tensors to OpenEnv observations and handle logging.
- **[`examples/grpo_blackjack/blackjack.yaml`](https://github.com/huggingface/OpenEnv/blob/main/examples/grpo_blackjack/blackjack.yaml)**: Training configuration defining reward shaping, policy hyperparameters, and rollout length.
- **[`README.md`](https://github.com/huggingface/OpenEnv/blob/main/README.md)** (line 337): Official torchforge integration documentation and entry point.

## Summary

- **Standardized Interface**: OpenEnv exposes containerized environments as Gymnasium-style interfaces via `EnvClient`, enabling direct integration with torchforge's `GRPOTrainer`.
- **Type Safety**: The client-server architecture uses Pydantic models for validation of observations, rewards, and done flags transmitted over WebSocket.
- **Infrastructure Abstraction**: Container providers in [`providers.py`](https://github.com/huggingface/OpenEnv/blob/main/providers.py) allow seamless swapping between Docker, Kubernetes, and cloud runtimes without modifying training code.
- **Reference Implementation**: The `examples/grpo_blackjack/` directory provides complete utilities for tensor conversion and GRPO loop wiring.

## Frequently Asked Questions

### How does torchforge handle the async nature of OpenEnv clients?

torchforge expects an async entry point for training loops. OpenEnv's `EnvClient` is async by default (see [`src/openenv/core/env_client.py`](https://github.com/huggingface/OpenEnv/blob/main/src/openenv/core/env_client.py) lines 65-68), making it compatible with torchforge's event-driven architecture. For synchronous trainers, wrap the client with `.sync()` to block on WebSocket calls.

### Can I use Kubernetes instead of Docker for the environment containers?

Yes. The `ProvisionerConfig` accepts a `provider` parameter that maps to implementations in [`src/openenv/core/containers/runtime/providers.py`](https://github.com/huggingface/OpenEnv/blob/main/src/openenv/core/containers/runtime/providers.py) (lines 10-16). Change `provider="docker"` to `provider="kubernetes"` to launch environments on a cluster without modifying the trainer logic.

### What is the MCP layer and when should I disable it?

MCP (Model Context Protocol) is an optional communication layer for multi-process setups. According to [`docs/source/tutorials/mcp-environment.md`](https://github.com/huggingface/OpenEnv/blob/main/docs/source/tutorials/mcp-environment.md) (line 46), you can bypass MCP when the trainer runs in the same process as the environment, which is the typical configuration for torchforge workflows.

### How are actions and observations typed between torchforge and OpenEnv?

OpenEnv environments define Pydantic models for actions and observations in the server implementation ([`mcp_environment.py`](https://github.com/huggingface/OpenEnv/blob/main/mcp_environment.py)). The `EnvClient` parses raw WebSocket JSON into these Python objects before returning them to torchforge, which then converts them to tensors via utility functions in [`examples/grpo_blackjack/grpo_utils.py`](https://github.com/huggingface/OpenEnv/blob/main/examples/grpo_blackjack/grpo_utils.py).