Configuring Delayed Rewards for Trajectory-Based Scoring in RL Training with OpenEnv

OpenEnv computes delayed rewards by accumulating trajectory data in rubrics and emitting a final score only when done=True, using the TrajectoryRubric class in src/openenv/core/rubrics/trajectory.py.

When training reinforcement learning agents on tasks where success is only determinable at episode completion, configuring delayed rewards for trajectory-based scoring becomes essential. The Hugging Face huggingface/OpenEnv library provides a modular rubric system that accumulates every (action, observation) pair throughout an episode and calculates a final scalar reward based on the complete trajectory. This approach avoids noisy step-wise signals and enables credit assignment based on overall behavioral quality.

How Trajectory-Based Scoring Works in OpenEnv

Trajectory Accumulation Mechanism

In src/openenv/core/rubrics/trajectory.py, the base TrajectoryRubric class maintains a private field self._trajectory: List[Tuple[Any, Any]] that stores every step's action and observation. Each call to step(action, observation, done) appends the pair to this internal list, building a complete history of the episode.

Delayed Reward Calculation

The rubric returns intermediate rewards (typically 0.0) during the episode. Only when done=True does the rubric invoke the abstract method score_trajectory(self._trajectory), which returns a scalar in [0, 1] representing the final delayed reward. This design ensures that the RL agent receives feedback based on the entire trajectory rather than individual transitions.

Configuring Trajectory Rubrics for Delayed Rewards

Using the Built-In ExponentialDiscountingTrajectoryRubric

OpenEnv ships with concrete implementations like ExponentialDiscountingTrajectoryRubric that apply time-discounting to trajectory steps. This class utilizes helper methods such as compute_step_rewards and exponential_discount defined in the base class to weight recent steps more heavily.

Implementing Custom Scoring Logic

For task-specific delayed rewards, subclass TrajectoryRubric and implement the score_trajectory method. The method receives the complete list of (action, observation) tuples and returns a float. Common implementations include binary success/failure indicators, success-rate calculations, or custom heuristic functions.

Episode Reset and State Management

The rubric integrates with the environment lifecycle through reset() calls. As implemented in src/openenv/core/env_server/interfaces.py, the reset hook clears the stored trajectory via rubric.reset(), preparing a fresh accumulator for the next episode. The state_dict() and load_state_dict() methods serialize rubric configuration (not the trajectory itself) for checkpointing.

Practical Code Examples

Example 1: Using ExponentialDiscountingTrajectoryRubric

from openenv.core.generic_client import GenericEnvClient
from openenv.core.rubrics.trajectory import ExponentialDiscountingTrajectoryRubric

client = GenericEnvClient(
    env_name="grid_world_env",
    rubric=ExponentialDiscountingTrajectoryRubric(
        discount=0.9,
        max_score=1.0,
    ),
)

obs = client.reset()
done = False
while not done:
    action = client.action_space.sample()
    obs, reward, done, info = client.step(action)

print("Episode finished – delayed reward:", reward)

Example 2: Custom Success Rate Trajectory Rubric

from typing import List, Tuple, Any
from openenv.core.rubrics.trajectory import TrajectoryRubric

class SuccessRateTrajectoryRubric(TrajectoryRubric):
    """Reward is the fraction of steps where observation['success']==True."""

    def score_trajectory(self, trajectory: List[Tuple[Any, Any]]) -> float:
        if not trajectory:
            return 0.0
        successes = sum(1 for _, obs in trajectory if obs.get("success"))
        return successes / len(trajectory)

client = GenericEnvClient(
    env_name="textarena_env",
    rubric=SuccessRateTrajectoryRubric(),
)

obs = client.reset()
done = False
while not done:
    action = client.action_space.sample()
    obs, reward, done, info = client.step(action)

print("Delayed success-rate reward:", reward)

Example 3: Explicit Reset Handling

client = GenericEnvClient(
    env_name="chess_env",
    rubric=ExponentialDiscountingTrajectoryRubric(),
)

# First episode

obs = client.reset()

# ... run steps ...

client.reset()  # triggers rubric.reset() and clears the trajectory

Key Source Files and Architecture

Summary

  • Trajectory rubrics accumulate (action, observation) pairs in self._trajectory throughout episodes.
  • Delayed rewards are computed by score_trajectory() only when done=True, returning a scalar in [0, 1].
  • Use ExponentialDiscountingTrajectoryRubric for built-in discounting or subclass TrajectoryRubric for custom logic.
  • Call reset() to clear trajectories between episodes, automatically handled via environment interfaces.
  • Configuration persists through state_dict()/load_state_dict() for checkpointing.

Frequently Asked Questions

When should I use delayed rewards instead of step-wise rewards?

Use delayed rewards when task success can only be determined after completing a full sequence, such as code generation, puzzle solving, or game completion, where intermediate actions have ambiguous value. OpenEnv's trajectory rubrics defer all credit assignment until the final step.

How does the trajectory rubric handle extremely long episodes?

The rubric stores all (action, observation) tuples in memory via self._trajectory. For very long episodes, consider implementing custom truncation logic in your score_trajectory method or using memory-efficient observation representations to prevent excessive memory consumption.

Can I combine trajectory-based scoring with other reward mechanisms?

Yes, though the base TrajectoryRubric is designed to return the final computed score. You may subclass and override the step() method to blend intermediate signals with the delayed final reward, or implement composition patterns that aggregate multiple rubric outputs.

Where is the trajectory data stored between steps?

The trajectory accumulates in the private attribute self._trajectory within the rubric instance, which persists in memory on the environment server until reset() is called or the episode terminates, as managed by the interfaces in src/openenv/core/env_server/interfaces.py.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →