Implementing Reward Functions with Domain-Specific Knowledge in OpenEnv Environments

OpenEnv encapsulates domain-specific reward logic inside composable Rubric objects that execute within environment servers, ensuring all reward signals originate from the environment and surface to training loops through StepResult payloads.

OpenEnv provides a structured framework for reinforcement learning environments where reward computation remains tightly coupled to domain expertise. By implementing reward functions with domain-specific knowledge inside OpenEnv environments, practitioners can encode complex evaluation criteria—from numeric tolerance checks to LLM-based judges—while maintaining clean separation between environment dynamics and training orchestration. The architecture centers on lightweight Rubric subclasses that transform (action, observation) pairs into scalar rewards.

The Rubric Abstraction

OpenEnv structures reward generation around Rubrics, lightweight objects defined in src/openenv/core/rubrics/base.py. The abstract Rubric class specifies a forward method where you implement domain-specific logic, and provides a callable interface that automatically executes pre‑forward and post‑forward hooks.

When subclassing Rubric, you override forward(action, observation) to return a float. The base class handles hook execution, making Rubrics suitable for both simple numeric checks and complex stateful evaluation:


# src/openenv/core/rubrics/base.py

from openenv.core.rubrics.base import Rubric

class NumericToleranceRubric(Rubric):
    """Reward 1.0 if prediction is within tol of target, else 0.0."""
    def __init__(self, target: float, tol: float = 0.01):
        super().__init__()
        self.target = target
        self.tol = tol

    def forward(self, action, observation) -> float:
        try:
            pred = float(action)
        except ValueError:
            return 0.0
        return 1.0 if abs(pred - self.target) <= self.tol else 0.0

Environment-Side Reward Generation

OpenEnv enforces a strict invariant: all reward signals must originate inside the environment. When an environment step executes, the server packs the result into a StepResult object that includes the computed reward. The MCP client (src/openenv/core/mcp_client.py) forwards this payload to the harness, which extracts the reward via _tool_result_reward without synthesizing new values.

The FinQA environment demonstrates this pattern in envs/finqa_env/server/finqa_environment.py. It imports domain-specific logic from envs/finqa_env/server/rewards.py, where compute_reward parses LaTeX-style boxed answers, normalizes percentages and fractions, and applies relative and absolute tolerance checks:


# envs/finqa_env/server/finqa_environment.py

from envs.finqa_env.server.rewards import compute_reward

class FinQAEnvironment:
    async def step(self, action):
        # action is the model's answer string

        ground_truth = self.current_question["answer"]
        reward = compute_reward(predicted=action, ground_truth=ground_truth)
        
        return {
            "observation": self._make_observation(),
            "reward": reward,  # Attached to StepResult payload

            "done": self._is_terminal(),
        }

Harness Integration and Reward Resolution

The harness (src/openenv/core/harness/__init__.py) finalizes rollout rewards by forwarding environment-produced values only. The function _resolve_env_reward guarantees that rewards are never manufactured by the orchestration layer, while _tool_result_reward extracts the scalar from the StepResult metadata:


# src/openenv/core/harness/__init__.py

def _tool_result_reward(tool_result):
    reward = tool_result.metadata.get("reward")
    if reward is None and isinstance(tool_result.data, dict):
        reward = tool_result.data.get("reward")
    return float(reward) if reward is not None else None

This design ensures that domain-specific logic remains traceable to the environment code itself, preventing reward hacking at the orchestration layer.

Composing Complex Reward Logic

Because Rubrics are standard Python objects, you can nest and compose them using utilities in src/openenv/core/rubrics/containers.py. This enables sophisticated evaluation criteria:

Composition allows you to combine multiple domain experts—for example, applying a numeric tolerance check first, then feeding borderline cases to an LLM judge.

Step-by-Step Implementation Workflow

To implement domain-specific rewards in your OpenEnv environment:

  1. Define the Rubric: Create a subclass of Rubric in your environment server or reuse a helper function that returns a float.
  2. Instantiate: Initialize the rubric inside your environment class constructor.
  3. Execute: Call the rubric's forward method (or helper function) during step(), passing the action and current observation.
  4. Payload: Store the resulting float in the StepResult return dictionary under the "reward" key.
  5. Consume: Use EnvClient from src/openenv/core/env_client.py to interact with the environment; the harness automatically surfaces the reward to your training loop.

# Client-side usage example

from openenv.core.env_client import EnvClient
import asyncio

async def main():
    client = await EnvClient.create("finqa")
    await client.reset()
    prediction = r"\boxed{42}"
    step = await client.step(prediction)
    print(f"Reward: {step.reward}")  # Value computed by environment

asyncio.run(main())

Summary

  • OpenEnv uses Rubric objects defined in src/openenv/core/rubrics/base.py to encapsulate reward logic with pre/post hooks.
  • Domain-specific rewards must originate inside the environment server, as enforced by _resolve_env_reward in the harness.
  • The FinQA environment (envs/finqa_env/server/rewards.py) demonstrates parsing LaTeX answers and applying numeric tolerance checks.
  • Rubric composition via containers.py supports complex criteria including trajectory discounting and LLM judges.
  • The MCP client forwards StepResult payloads containing rewards, which the harness extracts via _tool_result_reward without modification.

Frequently Asked Questions

How does OpenEnv ensure that rewards originate from the environment?

OpenEnv enforces this invariant through the harness function _resolve_env_reward in src/openenv/core/harness/__init__.py, which only forwards rewards found in the StepResult metadata. The harness never synthesizes reward values, ensuring that all signals are traceable to the environment's domain-specific implementation.

Can I combine multiple reward criteria in a single environment?

Yes. You can compose multiple Rubric instances using the container utilities in src/openenv/core/rubrics/containers.py. This allows you to aggregate scores from different domain experts—for example, combining a numeric tolerance check with an LLM judge—or apply trajectory-based discounting via src/openenv/core/rubrics/trajectory.py.

What is the difference between a Rubric and a raw reward function?

A Rubric is a formal subclass of the base class in src/openenv/core/rubrics/base.py that implements a forward method and supports pre/post hooks. A raw reward function is any callable that returns a float. You can wrap raw functions in Rubric subclasses to gain hook support, or use them directly inside your environment's step method as shown in the FinQA example.

How do I implement a time-dependent or trajectory-based reward?

Use the TrajectoryRubric class from src/openenv/core/rubrics/trajectory.py. This rubric maintains state across steps and can apply discount factors to rewards based on their position in the episode trajectory, enabling implementation of return-based objectives within the OpenEnv framework.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →