Implementing Reward Functions with Domain-Specific Knowledge in OpenEnv Environments
OpenEnv encapsulates domain-specific reward logic inside composable Rubric objects that execute within environment servers, ensuring all reward signals originate from the environment and surface to training loops through StepResult payloads.
OpenEnv provides a structured framework for reinforcement learning environments where reward computation remains tightly coupled to domain expertise. By implementing reward functions with domain-specific knowledge inside OpenEnv environments, practitioners can encode complex evaluation criteria—from numeric tolerance checks to LLM-based judges—while maintaining clean separation between environment dynamics and training orchestration. The architecture centers on lightweight Rubric subclasses that transform (action, observation) pairs into scalar rewards.
The Rubric Abstraction
OpenEnv structures reward generation around Rubrics, lightweight objects defined in src/openenv/core/rubrics/base.py. The abstract Rubric class specifies a forward method where you implement domain-specific logic, and provides a callable interface that automatically executes pre‑forward and post‑forward hooks.
When subclassing Rubric, you override forward(action, observation) to return a float. The base class handles hook execution, making Rubrics suitable for both simple numeric checks and complex stateful evaluation:
# src/openenv/core/rubrics/base.py
from openenv.core.rubrics.base import Rubric
class NumericToleranceRubric(Rubric):
"""Reward 1.0 if prediction is within tol of target, else 0.0."""
def __init__(self, target: float, tol: float = 0.01):
super().__init__()
self.target = target
self.tol = tol
def forward(self, action, observation) -> float:
try:
pred = float(action)
except ValueError:
return 0.0
return 1.0 if abs(pred - self.target) <= self.tol else 0.0
Environment-Side Reward Generation
OpenEnv enforces a strict invariant: all reward signals must originate inside the environment. When an environment step executes, the server packs the result into a StepResult object that includes the computed reward. The MCP client (src/openenv/core/mcp_client.py) forwards this payload to the harness, which extracts the reward via _tool_result_reward without synthesizing new values.
The FinQA environment demonstrates this pattern in envs/finqa_env/server/finqa_environment.py. It imports domain-specific logic from envs/finqa_env/server/rewards.py, where compute_reward parses LaTeX-style boxed answers, normalizes percentages and fractions, and applies relative and absolute tolerance checks:
# envs/finqa_env/server/finqa_environment.py
from envs.finqa_env.server.rewards import compute_reward
class FinQAEnvironment:
async def step(self, action):
# action is the model's answer string
ground_truth = self.current_question["answer"]
reward = compute_reward(predicted=action, ground_truth=ground_truth)
return {
"observation": self._make_observation(),
"reward": reward, # Attached to StepResult payload
"done": self._is_terminal(),
}
Harness Integration and Reward Resolution
The harness (src/openenv/core/harness/__init__.py) finalizes rollout rewards by forwarding environment-produced values only. The function _resolve_env_reward guarantees that rewards are never manufactured by the orchestration layer, while _tool_result_reward extracts the scalar from the StepResult metadata:
# src/openenv/core/harness/__init__.py
def _tool_result_reward(tool_result):
reward = tool_result.metadata.get("reward")
if reward is None and isinstance(tool_result.data, dict):
reward = tool_result.data.get("reward")
return float(reward) if reward is not None else None
This design ensures that domain-specific logic remains traceable to the environment code itself, preventing reward hacking at the orchestration layer.
Composing Complex Reward Logic
Because Rubrics are standard Python objects, you can nest and compose them using utilities in src/openenv/core/rubrics/containers.py. This enables sophisticated evaluation criteria:
- Trajectory-based discounting: The
TrajectoryRubricclass insrc/openenv/core/rubrics/trajectory.pyapplies time-step discounts across episode histories. - LLM-as-judge: The
LLMJudgeRubricinsrc/openenv/core/rubrics/llm_judge.pyuses language models to evaluate free-form text generations against rubric criteria.
Composition allows you to combine multiple domain experts—for example, applying a numeric tolerance check first, then feeding borderline cases to an LLM judge.
Step-by-Step Implementation Workflow
To implement domain-specific rewards in your OpenEnv environment:
- Define the Rubric: Create a subclass of
Rubricin your environment server or reuse a helper function that returns a float. - Instantiate: Initialize the rubric inside your environment class constructor.
- Execute: Call the rubric's
forwardmethod (or helper function) duringstep(), passing the action and current observation. - Payload: Store the resulting float in the
StepResultreturn dictionary under the"reward"key. - Consume: Use
EnvClientfromsrc/openenv/core/env_client.pyto interact with the environment; the harness automatically surfaces the reward to your training loop.
# Client-side usage example
from openenv.core.env_client import EnvClient
import asyncio
async def main():
client = await EnvClient.create("finqa")
await client.reset()
prediction = r"\boxed{42}"
step = await client.step(prediction)
print(f"Reward: {step.reward}") # Value computed by environment
asyncio.run(main())
Summary
- OpenEnv uses Rubric objects defined in
src/openenv/core/rubrics/base.pyto encapsulate reward logic with pre/post hooks. - Domain-specific rewards must originate inside the environment server, as enforced by
_resolve_env_rewardin the harness. - The FinQA environment (
envs/finqa_env/server/rewards.py) demonstrates parsing LaTeX answers and applying numeric tolerance checks. - Rubric composition via
containers.pysupports complex criteria including trajectory discounting and LLM judges. - The MCP client forwards
StepResultpayloads containing rewards, which the harness extracts via_tool_result_rewardwithout modification.
Frequently Asked Questions
How does OpenEnv ensure that rewards originate from the environment?
OpenEnv enforces this invariant through the harness function _resolve_env_reward in src/openenv/core/harness/__init__.py, which only forwards rewards found in the StepResult metadata. The harness never synthesizes reward values, ensuring that all signals are traceable to the environment's domain-specific implementation.
Can I combine multiple reward criteria in a single environment?
Yes. You can compose multiple Rubric instances using the container utilities in src/openenv/core/rubrics/containers.py. This allows you to aggregate scores from different domain experts—for example, combining a numeric tolerance check with an LLM judge—or apply trajectory-based discounting via src/openenv/core/rubrics/trajectory.py.
What is the difference between a Rubric and a raw reward function?
A Rubric is a formal subclass of the base class in src/openenv/core/rubrics/base.py that implements a forward method and supports pre/post hooks. A raw reward function is any callable that returns a float. You can wrap raw functions in Rubric subclasses to gain hook support, or use them directly inside your environment's step method as shown in the FinQA example.
How do I implement a time-dependent or trajectory-based reward?
Use the TrajectoryRubric class from src/openenv/core/rubrics/trajectory.py. This rubric maintains state across steps and can apply discount factors to rewards based on their position in the episode trajectory, enabling implementation of return-based objectives within the OpenEnv framework.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →