Configuring Delayed Rewards for Trajectory-Based Scoring in RL Training with OpenEnv
OpenEnv computes delayed rewards by accumulating trajectory data in rubrics and emitting a final score only when done=True, using the TrajectoryRubric class in src/openenv/core/rubrics/trajectory.py.
When training reinforcement learning agents on tasks where success is only determinable at episode completion, configuring delayed rewards for trajectory-based scoring becomes essential. The Hugging Face huggingface/OpenEnv library provides a modular rubric system that accumulates every (action, observation) pair throughout an episode and calculates a final scalar reward based on the complete trajectory. This approach avoids noisy step-wise signals and enables credit assignment based on overall behavioral quality.
How Trajectory-Based Scoring Works in OpenEnv
Trajectory Accumulation Mechanism
In src/openenv/core/rubrics/trajectory.py, the base TrajectoryRubric class maintains a private field self._trajectory: List[Tuple[Any, Any]] that stores every step's action and observation. Each call to step(action, observation, done) appends the pair to this internal list, building a complete history of the episode.
Delayed Reward Calculation
The rubric returns intermediate rewards (typically 0.0) during the episode. Only when done=True does the rubric invoke the abstract method score_trajectory(self._trajectory), which returns a scalar in [0, 1] representing the final delayed reward. This design ensures that the RL agent receives feedback based on the entire trajectory rather than individual transitions.
Configuring Trajectory Rubrics for Delayed Rewards
Using the Built-In ExponentialDiscountingTrajectoryRubric
OpenEnv ships with concrete implementations like ExponentialDiscountingTrajectoryRubric that apply time-discounting to trajectory steps. This class utilizes helper methods such as compute_step_rewards and exponential_discount defined in the base class to weight recent steps more heavily.
Implementing Custom Scoring Logic
For task-specific delayed rewards, subclass TrajectoryRubric and implement the score_trajectory method. The method receives the complete list of (action, observation) tuples and returns a float. Common implementations include binary success/failure indicators, success-rate calculations, or custom heuristic functions.
Episode Reset and State Management
The rubric integrates with the environment lifecycle through reset() calls. As implemented in src/openenv/core/env_server/interfaces.py, the reset hook clears the stored trajectory via rubric.reset(), preparing a fresh accumulator for the next episode. The state_dict() and load_state_dict() methods serialize rubric configuration (not the trajectory itself) for checkpointing.
Practical Code Examples
Example 1: Using ExponentialDiscountingTrajectoryRubric
from openenv.core.generic_client import GenericEnvClient
from openenv.core.rubrics.trajectory import ExponentialDiscountingTrajectoryRubric
client = GenericEnvClient(
env_name="grid_world_env",
rubric=ExponentialDiscountingTrajectoryRubric(
discount=0.9,
max_score=1.0,
),
)
obs = client.reset()
done = False
while not done:
action = client.action_space.sample()
obs, reward, done, info = client.step(action)
print("Episode finished – delayed reward:", reward)
Example 2: Custom Success Rate Trajectory Rubric
from typing import List, Tuple, Any
from openenv.core.rubrics.trajectory import TrajectoryRubric
class SuccessRateTrajectoryRubric(TrajectoryRubric):
"""Reward is the fraction of steps where observation['success']==True."""
def score_trajectory(self, trajectory: List[Tuple[Any, Any]]) -> float:
if not trajectory:
return 0.0
successes = sum(1 for _, obs in trajectory if obs.get("success"))
return successes / len(trajectory)
client = GenericEnvClient(
env_name="textarena_env",
rubric=SuccessRateTrajectoryRubric(),
)
obs = client.reset()
done = False
while not done:
action = client.action_space.sample()
obs, reward, done, info = client.step(action)
print("Delayed success-rate reward:", reward)
Example 3: Explicit Reset Handling
client = GenericEnvClient(
env_name="chess_env",
rubric=ExponentialDiscountingTrajectoryRubric(),
)
# First episode
obs = client.reset()
# ... run steps ...
client.reset() # triggers rubric.reset() and clears the trajectory
Key Source Files and Architecture
src/openenv/core/rubrics/trajectory.py– Core implementation ofTrajectoryRubricand subclasses includingExponentialDiscountingTrajectoryRubric.src/openenv/core/env_server/interfaces.py– Provides the reset hook that clears rubric trajectories between episodes.src/openenv/core/generic_client.py– High-level client that wires environments with rubrics.tests/core/test_rubrics/test_trajectory_rubric.py– Unit tests verifying trajectory accumulation, scoring, and serialization.tests/core/test_rubrics/test_environment_integration.py– Integration tests demonstrating rubric behavior when attached to environments.
Summary
- Trajectory rubrics accumulate
(action, observation)pairs inself._trajectorythroughout episodes. - Delayed rewards are computed by
score_trajectory()only whendone=True, returning a scalar in[0, 1]. - Use
ExponentialDiscountingTrajectoryRubricfor built-in discounting or subclassTrajectoryRubricfor custom logic. - Call
reset()to clear trajectories between episodes, automatically handled via environment interfaces. - Configuration persists through
state_dict()/load_state_dict()for checkpointing.
Frequently Asked Questions
When should I use delayed rewards instead of step-wise rewards?
Use delayed rewards when task success can only be determined after completing a full sequence, such as code generation, puzzle solving, or game completion, where intermediate actions have ambiguous value. OpenEnv's trajectory rubrics defer all credit assignment until the final step.
How does the trajectory rubric handle extremely long episodes?
The rubric stores all (action, observation) tuples in memory via self._trajectory. For very long episodes, consider implementing custom truncation logic in your score_trajectory method or using memory-efficient observation representations to prevent excessive memory consumption.
Can I combine trajectory-based scoring with other reward mechanisms?
Yes, though the base TrajectoryRubric is designed to return the final computed score. You may subclass and override the step() method to blend intermediate signals with the delayed final reward, or implement composition patterns that aggregate multiple rubric outputs.
Where is the trajectory data stored between steps?
The trajectory accumulates in the private attribute self._trajectory within the rubric instance, which persists in memory on the environment server until reset() is called or the episode terminates, as managed by the interfaces in src/openenv/core/env_server/interfaces.py.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →