# Configuring Delayed Rewards for Trajectory-Based Scoring in RL Training with OpenEnv

> Learn to configure delayed rewards for trajectory-based scoring in RL training with OpenEnv. Discover how to accumulate trajectory data and emit final scores effectively.

- Repository: [Hugging Face/OpenEnv](https://github.com/huggingface/OpenEnv)
- Tags: how-to-guide
- Published: 2026-06-14

---

**OpenEnv computes delayed rewards by accumulating trajectory data in rubrics and emitting a final score only when `done=True`, using the `TrajectoryRubric` class in [`src/openenv/core/rubrics/trajectory.py`](https://github.com/huggingface/OpenEnv/blob/main/src/openenv/core/rubrics/trajectory.py).**

When training reinforcement learning agents on tasks where success is only determinable at episode completion, configuring delayed rewards for trajectory-based scoring becomes essential. The Hugging Face `huggingface/OpenEnv` library provides a modular rubric system that accumulates every `(action, observation)` pair throughout an episode and calculates a final scalar reward based on the complete trajectory. This approach avoids noisy step-wise signals and enables credit assignment based on overall behavioral quality.

## How Trajectory-Based Scoring Works in OpenEnv

### Trajectory Accumulation Mechanism

In [`src/openenv/core/rubrics/trajectory.py`](https://github.com/huggingface/OpenEnv/blob/main/src/openenv/core/rubrics/trajectory.py), the base `TrajectoryRubric` class maintains a private field `self._trajectory: List[Tuple[Any, Any]]` that stores every step's action and observation. Each call to `step(action, observation, done)` appends the pair to this internal list, building a complete history of the episode.

### Delayed Reward Calculation

The rubric returns intermediate rewards (typically `0.0`) during the episode. Only when `done=True` does the rubric invoke the abstract method `score_trajectory(self._trajectory)`, which returns a scalar in `[0, 1]` representing the final delayed reward. This design ensures that the RL agent receives feedback based on the entire trajectory rather than individual transitions.

## Configuring Trajectory Rubrics for Delayed Rewards

### Using the Built-In ExponentialDiscountingTrajectoryRubric

OpenEnv ships with concrete implementations like `ExponentialDiscountingTrajectoryRubric` that apply time-discounting to trajectory steps. This class utilizes helper methods such as `compute_step_rewards` and `exponential_discount` defined in the base class to weight recent steps more heavily.

### Implementing Custom Scoring Logic

For task-specific delayed rewards, subclass `TrajectoryRubric` and implement the `score_trajectory` method. The method receives the complete list of `(action, observation)` tuples and returns a float. Common implementations include binary success/failure indicators, success-rate calculations, or custom heuristic functions.

### Episode Reset and State Management

The rubric integrates with the environment lifecycle through `reset()` calls. As implemented in [`src/openenv/core/env_server/interfaces.py`](https://github.com/huggingface/OpenEnv/blob/main/src/openenv/core/env_server/interfaces.py), the reset hook clears the stored trajectory via `rubric.reset()`, preparing a fresh accumulator for the next episode. The `state_dict()` and `load_state_dict()` methods serialize rubric configuration (not the trajectory itself) for checkpointing.

## Practical Code Examples

### Example 1: Using ExponentialDiscountingTrajectoryRubric

```python
from openenv.core.generic_client import GenericEnvClient
from openenv.core.rubrics.trajectory import ExponentialDiscountingTrajectoryRubric

client = GenericEnvClient(
    env_name="grid_world_env",
    rubric=ExponentialDiscountingTrajectoryRubric(
        discount=0.9,
        max_score=1.0,
    ),
)

obs = client.reset()
done = False
while not done:
    action = client.action_space.sample()
    obs, reward, done, info = client.step(action)

print("Episode finished – delayed reward:", reward)

```

### Example 2: Custom Success Rate Trajectory Rubric

```python
from typing import List, Tuple, Any
from openenv.core.rubrics.trajectory import TrajectoryRubric

class SuccessRateTrajectoryRubric(TrajectoryRubric):
    """Reward is the fraction of steps where observation['success']==True."""

    def score_trajectory(self, trajectory: List[Tuple[Any, Any]]) -> float:
        if not trajectory:
            return 0.0
        successes = sum(1 for _, obs in trajectory if obs.get("success"))
        return successes / len(trajectory)

client = GenericEnvClient(
    env_name="textarena_env",
    rubric=SuccessRateTrajectoryRubric(),
)

obs = client.reset()
done = False
while not done:
    action = client.action_space.sample()
    obs, reward, done, info = client.step(action)

print("Delayed success-rate reward:", reward)

```

### Example 3: Explicit Reset Handling

```python
client = GenericEnvClient(
    env_name="chess_env",
    rubric=ExponentialDiscountingTrajectoryRubric(),
)

# First episode

obs = client.reset()

# ... run steps ...

client.reset()  # triggers rubric.reset() and clears the trajectory

```

## Key Source Files and Architecture

- **[`src/openenv/core/rubrics/trajectory.py`](https://github.com/huggingface/OpenEnv/blob/main/src/openenv/core/rubrics/trajectory.py)** – Core implementation of `TrajectoryRubric` and subclasses including `ExponentialDiscountingTrajectoryRubric`.
- **[`src/openenv/core/env_server/interfaces.py`](https://github.com/huggingface/OpenEnv/blob/main/src/openenv/core/env_server/interfaces.py)** – Provides the reset hook that clears rubric trajectories between episodes.
- **[`src/openenv/core/generic_client.py`](https://github.com/huggingface/OpenEnv/blob/main/src/openenv/core/generic_client.py)** – High-level client that wires environments with rubrics.
- **[`tests/core/test_rubrics/test_trajectory_rubric.py`](https://github.com/huggingface/OpenEnv/blob/main/tests/core/test_rubrics/test_trajectory_rubric.py)** – Unit tests verifying trajectory accumulation, scoring, and serialization.
- **[`tests/core/test_rubrics/test_environment_integration.py`](https://github.com/huggingface/OpenEnv/blob/main/tests/core/test_rubrics/test_environment_integration.py)** – Integration tests demonstrating rubric behavior when attached to environments.

## Summary

- Trajectory rubrics accumulate `(action, observation)` pairs in `self._trajectory` throughout episodes.
- **Delayed rewards** are computed by `score_trajectory()` only when `done=True`, returning a scalar in `[0, 1]`.
- Use `ExponentialDiscountingTrajectoryRubric` for built-in discounting or subclass `TrajectoryRubric` for custom logic.
- Call `reset()` to clear trajectories between episodes, automatically handled via environment interfaces.
- Configuration persists through `state_dict()`/`load_state_dict()` for checkpointing.

## Frequently Asked Questions

### When should I use delayed rewards instead of step-wise rewards?

Use delayed rewards when task success can only be determined after completing a full sequence, such as code generation, puzzle solving, or game completion, where intermediate actions have ambiguous value. OpenEnv's trajectory rubrics defer all credit assignment until the final step.

### How does the trajectory rubric handle extremely long episodes?

The rubric stores all `(action, observation)` tuples in memory via `self._trajectory`. For very long episodes, consider implementing custom truncation logic in your `score_trajectory` method or using memory-efficient observation representations to prevent excessive memory consumption.

### Can I combine trajectory-based scoring with other reward mechanisms?

Yes, though the base `TrajectoryRubric` is designed to return the final computed score. You may subclass and override the `step()` method to blend intermediate signals with the delayed final reward, or implement composition patterns that aggregate multiple rubric outputs.

### Where is the trajectory data stored between steps?

The trajectory accumulates in the private attribute `self._trajectory` within the rubric instance, which persists in memory on the environment server until `reset()` is called or the episode terminates, as managed by the interfaces in [`src/openenv/core/env_server/interfaces.py`](https://github.com/huggingface/OpenEnv/blob/main/src/openenv/core/env_server/interfaces.py).