# Implementing Reward Functions with Domain-Specific Knowledge in OpenEnv Environments

> Learn to implement reward functions with domain-specific knowledge in OpenEnv using Rubric objects. Ensure all reward signals come from the environment and are available through StepResult payloads.

- Repository: [Hugging Face/OpenEnv](https://github.com/huggingface/OpenEnv)
- Tags: how-to-guide
- Published: 2026-06-14

---

**OpenEnv encapsulates domain-specific reward logic inside composable Rubric objects that execute within environment servers, ensuring all reward signals originate from the environment and surface to training loops through StepResult payloads.**

OpenEnv provides a structured framework for reinforcement learning environments where reward computation remains tightly coupled to domain expertise. By implementing reward functions with domain-specific knowledge inside OpenEnv environments, practitioners can encode complex evaluation criteria—from numeric tolerance checks to LLM-based judges—while maintaining clean separation between environment dynamics and training orchestration. The architecture centers on lightweight **Rubric** subclasses that transform (action, observation) pairs into scalar rewards.

## The Rubric Abstraction

OpenEnv structures reward generation around **Rubrics**, lightweight objects defined in [`src/openenv/core/rubrics/base.py`](https://github.com/huggingface/OpenEnv/blob/main/src/openenv/core/rubrics/base.py). The abstract `Rubric` class specifies a `forward` method where you implement domain-specific logic, and provides a callable interface that automatically executes pre‑forward and post‑forward hooks.

When subclassing `Rubric`, you override `forward(action, observation)` to return a float. The base class handles hook execution, making Rubrics suitable for both simple numeric checks and complex stateful evaluation:

```python

# src/openenv/core/rubrics/base.py

from openenv.core.rubrics.base import Rubric

class NumericToleranceRubric(Rubric):
    """Reward 1.0 if prediction is within tol of target, else 0.0."""
    def __init__(self, target: float, tol: float = 0.01):
        super().__init__()
        self.target = target
        self.tol = tol

    def forward(self, action, observation) -> float:
        try:
            pred = float(action)
        except ValueError:
            return 0.0
        return 1.0 if abs(pred - self.target) <= self.tol else 0.0

```

## Environment-Side Reward Generation

OpenEnv enforces a strict invariant: **all reward signals must originate inside the environment**. When an environment step executes, the server packs the result into a `StepResult` object that includes the computed reward. The MCP client ([`src/openenv/core/mcp_client.py`](https://github.com/huggingface/OpenEnv/blob/main/src/openenv/core/mcp_client.py)) forwards this payload to the harness, which extracts the reward via `_tool_result_reward` without synthesizing new values.

The FinQA environment demonstrates this pattern in [`envs/finqa_env/server/finqa_environment.py`](https://github.com/huggingface/OpenEnv/blob/main/envs/finqa_env/server/finqa_environment.py). It imports domain-specific logic from [`envs/finqa_env/server/rewards.py`](https://github.com/huggingface/OpenEnv/blob/main/envs/finqa_env/server/rewards.py), where `compute_reward` parses LaTeX-style boxed answers, normalizes percentages and fractions, and applies relative and absolute tolerance checks:

```python

# envs/finqa_env/server/finqa_environment.py

from envs.finqa_env.server.rewards import compute_reward

class FinQAEnvironment:
    async def step(self, action):
        # action is the model's answer string

        ground_truth = self.current_question["answer"]
        reward = compute_reward(predicted=action, ground_truth=ground_truth)
        
        return {
            "observation": self._make_observation(),
            "reward": reward,  # Attached to StepResult payload

            "done": self._is_terminal(),
        }

```

## Harness Integration and Reward Resolution

The harness ([`src/openenv/core/harness/__init__.py`](https://github.com/huggingface/OpenEnv/blob/main/src/openenv/core/harness/__init__.py)) finalizes rollout rewards by forwarding environment-produced values only. The function `_resolve_env_reward` guarantees that rewards are never manufactured by the orchestration layer, while `_tool_result_reward` extracts the scalar from the `StepResult` metadata:

```python

# src/openenv/core/harness/__init__.py

def _tool_result_reward(tool_result):
    reward = tool_result.metadata.get("reward")
    if reward is None and isinstance(tool_result.data, dict):
        reward = tool_result.data.get("reward")
    return float(reward) if reward is not None else None

```

This design ensures that domain-specific logic remains traceable to the environment code itself, preventing reward hacking at the orchestration layer.

## Composing Complex Reward Logic

Because Rubrics are standard Python objects, you can nest and compose them using utilities in [`src/openenv/core/rubrics/containers.py`](https://github.com/huggingface/OpenEnv/blob/main/src/openenv/core/rubrics/containers.py). This enables sophisticated evaluation criteria:

- **Trajectory-based discounting**: The `TrajectoryRubric` class in [`src/openenv/core/rubrics/trajectory.py`](https://github.com/huggingface/OpenEnv/blob/main/src/openenv/core/rubrics/trajectory.py) applies time-step discounts across episode histories.
- **LLM-as-judge**: The `LLMJudgeRubric` in [`src/openenv/core/rubrics/llm_judge.py`](https://github.com/huggingface/OpenEnv/blob/main/src/openenv/core/rubrics/llm_judge.py) uses language models to evaluate free-form text generations against rubric criteria.

Composition allows you to combine multiple domain experts—for example, applying a numeric tolerance check first, then feeding borderline cases to an LLM judge.

## Step-by-Step Implementation Workflow

To implement domain-specific rewards in your OpenEnv environment:

1. **Define the Rubric**: Create a subclass of `Rubric` in your environment server or reuse a helper function that returns a float.
2. **Instantiate**: Initialize the rubric inside your environment class constructor.
3. **Execute**: Call the rubric's `forward` method (or helper function) during `step()`, passing the action and current observation.
4. **Payload**: Store the resulting float in the `StepResult` return dictionary under the `"reward"` key.
5. **Consume**: Use `EnvClient` from [`src/openenv/core/env_client.py`](https://github.com/huggingface/OpenEnv/blob/main/src/openenv/core/env_client.py) to interact with the environment; the harness automatically surfaces the reward to your training loop.

```python

# Client-side usage example

from openenv.core.env_client import EnvClient
import asyncio

async def main():
    client = await EnvClient.create("finqa")
    await client.reset()
    prediction = r"\boxed{42}"
    step = await client.step(prediction)
    print(f"Reward: {step.reward}")  # Value computed by environment

asyncio.run(main())

```

## Summary

- OpenEnv uses **Rubric** objects defined in [`src/openenv/core/rubrics/base.py`](https://github.com/huggingface/OpenEnv/blob/main/src/openenv/core/rubrics/base.py) to encapsulate reward logic with pre/post hooks.
- **Domain-specific rewards** must originate inside the environment server, as enforced by `_resolve_env_reward` in the harness.
- The **FinQA environment** ([`envs/finqa_env/server/rewards.py`](https://github.com/huggingface/OpenEnv/blob/main/envs/finqa_env/server/rewards.py)) demonstrates parsing LaTeX answers and applying numeric tolerance checks.
- **Rubric composition** via [`containers.py`](https://github.com/huggingface/OpenEnv/blob/main/containers.py) supports complex criteria including trajectory discounting and LLM judges.
- The **MCP client** forwards `StepResult` payloads containing rewards, which the harness extracts via `_tool_result_reward` without modification.

## Frequently Asked Questions

### How does OpenEnv ensure that rewards originate from the environment?

OpenEnv enforces this invariant through the harness function `_resolve_env_reward` in [`src/openenv/core/harness/__init__.py`](https://github.com/huggingface/OpenEnv/blob/main/src/openenv/core/harness/__init__.py), which only forwards rewards found in the `StepResult` metadata. The harness never synthesizes reward values, ensuring that all signals are traceable to the environment's domain-specific implementation.

### Can I combine multiple reward criteria in a single environment?

Yes. You can compose multiple Rubric instances using the container utilities in [`src/openenv/core/rubrics/containers.py`](https://github.com/huggingface/OpenEnv/blob/main/src/openenv/core/rubrics/containers.py). This allows you to aggregate scores from different domain experts—for example, combining a numeric tolerance check with an LLM judge—or apply trajectory-based discounting via [`src/openenv/core/rubrics/trajectory.py`](https://github.com/huggingface/OpenEnv/blob/main/src/openenv/core/rubrics/trajectory.py).

### What is the difference between a Rubric and a raw reward function?

A **Rubric** is a formal subclass of the base class in [`src/openenv/core/rubrics/base.py`](https://github.com/huggingface/OpenEnv/blob/main/src/openenv/core/rubrics/base.py) that implements a `forward` method and supports pre/post hooks. A raw reward function is any callable that returns a float. You can wrap raw functions in Rubric subclasses to gain hook support, or use them directly inside your environment's `step` method as shown in the FinQA example.

### How do I implement a time-dependent or trajectory-based reward?

Use the `TrajectoryRubric` class from [`src/openenv/core/rubrics/trajectory.py`](https://github.com/huggingface/OpenEnv/blob/main/src/openenv/core/rubrics/trajectory.py). This rubric maintains state across steps and can apply discount factors to rewards based on their position in the episode trajectory, enabling implementation of return-based objectives within the OpenEnv framework.