# How the OpenRisk Eval Framework Evaluates Agent Performance and Accuracy

> Learn how the OpenRisk Eval framework assesses agent performance and accuracy. It uses a data-flow DAG to capture, score predictions, and return structured evaluation results.

- Repository: [derisk-ai/openderisk](https://github.com/derisk-ai/openderisk)
- Tags: performance
- Published: 2026-02-28

---

**The OpenRisk evaluation framework quantifies agent performance and accuracy by orchestrating a data-flow DAG that captures predictions via `AgentOutputOperator`, scores them against ground-truth contexts using pluggable `EvaluationMetric` implementations, and returns structured `EvaluationResult` objects containing numeric scores, pass/fail flags, and latency metrics.**

The `derisk-ai/openderisk` repository provides a modular evaluation system designed to benchmark LLM-driven agents and retrievers. By combining abstract metric definitions with a directed acyclic graph (DAG) execution model, the framework enables precise measurement of both semantic relevance and task-specific correctness.

## Core Evaluation Abstractions

### The EvaluationMetric Contract

All accuracy measurements in OpenRisk derive from the abstract `EvaluationMetric` class defined in [`packages/derisk-core/src/derisk/core/interface/evaluation.py`](https://github.com/derisk-ai/openderisk/blob/main/packages/derisk-core/src/derisk/core/interface/evaluation.py). This contract requires a single synchronous method:

```python
def sync_compute(self, prediction, contexts, query=None) -> BaseEvaluationResult:

```

Concrete implementations override this method to define domain-specific scoring logic. For example, `RetrieverSimilarityMetric` computes cosine similarity between the agent's answer and ground-truth contexts, while `IntentMetric` parses JSON predictions to verify exact intent matches. Each metric returns a `BaseEvaluationResult` containing the raw score, a boolean `passing` flag, and optional feedback text.

### Result Containers and Registry

The framework uses two primary data containers also located in [`packages/derisk-core/src/derisk/core/interface/evaluation.py`](https://github.com/derisk-ai/openderisk/blob/main/packages/derisk-core/src/derisk/core/interface/evaluation.py):

- **`BaseEvaluationResult`**: Holds the numeric `score`, `prediction` text, `passing` boolean, and `contexts` list for a single metric computation.
- **`EvaluationResult`**: Wraps `BaseEvaluationResult` with additional metadata including the metric name, query string, raw dataset entry, and `prediction_cost` (latency).

The `MetricManage` class maintains a global registry mapping metric names to their classes. Developers register new metrics using `metric_manage.register_metric(MyMetric)`, enabling dynamic lookup during evaluation runs.

## The Agent Evaluation DAG

The evaluation pipeline is implemented as a data-flow DAG in [`packages/derisk-serve/src/derisk_serve/agent/evaluation/evaluation.py`](https://github.com/derisk-ai/openderisk/blob/main/packages/derisk-serve/src/derisk_serve/agent/evaluation/evaluation.py). Three specialized operators coordinate the workflow:

### AgentOutputOperator (Prediction Capture)

`AgentOutputOperator` is a `MapOperator` that invokes the target agent via `multi_agents.app_chat` (defined in [`packages/derisk-serve/src/derisk_serve/agent/agents/controller.py`](https://github.com/derisk-ai/openderisk/blob/main/packages/derisk-serve/src/derisk_serve/agent/agents/controller.py)). It captures the final answer and records execution time as `prediction_cost`. This operator transforms a raw query into an agent prediction ready for scoring.

### AgentEvaluatorOperator (Scoring)

`AgentEvaluatorOperator` acts as a `JoinOperator` that receives the original query, the agent's prediction, and the ground-truth contexts. For each selected metric, it asynchronously executes `metric.compute(prediction, contexts)` and assembles the outputs into `EvaluationResult` objects.

### AgentEvaluator (Orchestration)

The `AgentEvaluator` class (subclass of `Evaluator`) orchestrates the full pipeline:
1. **Dataset ingestion**: An `IteratorTrigger` yields rows containing queries and contexts.
2. **Query extraction**: A `MapOperator` extracts the query field specified by `query_key`.
3. **Agent invocation**: `AgentOutputOperator` generates predictions.
4. **Metric evaluation**: `AgentEvaluatorOperator` applies all registered metrics.
5. **Result aggregation**: Returns a `List[List[EvaluationResult]]` where outer indices represent dataset rows and inner indices represent individual metrics.

## Built-in Accuracy Metrics

OpenRisk ships with concrete metrics covering both relevance and correctness dimensions:

- **RetrieverSimilarityMetric**: Uses embeddings to calculate cosine similarity between predictions and contexts.
- **RetrieverHitRateMetric**: Binary score indicating whether the correct context appears in the top-k retrievals.
- **AnswerRelevancyMetric**: Assesses semantic relevance of the generated answer to the query.
- **AppLinkMetric & IntentMetric**: Parse JSON-encoded predictions to verify exact matches for `app_name` or `intent` fields against expected values.

These metrics are registered in [`packages/derisk-serve/src/derisk_serve/agent/evaluation/evaluation_metric.py`](https://github.com/derisk-ai/openderisk/blob/main/packages/derisk-serve/src/derisk_serve/agent/evaluation/evaluation_metric.py) and `packages/derisk-rag/src/derisk/rag/evaluation/`.

## Running an Evaluation

To evaluate agent performance against a dataset with the default similarity metric:

```python
import asyncio
from derisk.rag.evaluation import RetrieverSimilarityMetric
from derisk_serve.agent.evaluation.evaluation import AgentEvaluator, AgentOutputOperator

dataset = [
    {
        "query": "What does the OpenRisk API provide?",
        "contexts": [
            "OpenRisk offers risk-assessment APIs for financial models.",
        ],
    },
]

evaluator = AgentEvaluator(
    operator_cls=AgentOutputOperator,
    operator_kwargs={"app_code": "my_agent"},
)

results = asyncio.run(evaluator.evaluate(dataset))

# Inspect results: results[row_index][metric_index]

print(results[0][0].metric_name)   # "RetrieverSimilarityMetric"

print(results[0][0].score)         # Numeric similarity score

print(results[0][0].passing)       # True/False

print(results[0][0].prediction_cost)  # Latency in seconds

```

## Extending the Framework

Add custom accuracy checks by subclassing `EvaluationMetric` and registering with `MetricManage`:

```python
from derisk.core.interface.evaluation import EvaluationMetric, BaseEvaluationResult, metric_manage
import json

class IntentCorrectnessMetric(EvaluationMetric[str, str]):
    def sync_compute(self, prediction, contexts, query=None):
        try:
            parsed = json.loads(prediction)
            intent = parsed.get("intent")
            score = 1.0 if intent in contexts else 0.0
            passing = bool(score)
        except Exception:
            score = 0.0
            passing = False
            intent = None
        
        return BaseEvaluationResult(
            score=score, 
            prediction=intent, 
            passing=passing
        )

# Register for dynamic discovery

metric_manage.register_metric(IntentCorrectnessMetric)

# Use in evaluation

custom_evaluator = AgentEvaluator(
    operator_cls=AgentOutputOperator,
    operator_kwargs={"app_code": "my_agent"},
)
results = asyncio.run(custom_evaluator.evaluate(
    dataset, 
    metrics=[IntentCorrectnessMetric()]
))

```

## Summary

- The OpenRisk evaluation framework uses a DAG architecture to assess agent performance and accuracy through composable operators.
- **`EvaluationMetric`** defines the scoring contract via `sync_compute`, while **`MetricManage`** enables dynamic metric registration.
- **`AgentOutputOperator`** captures predictions and latency (`prediction_cost`) by calling `multi_agents.app_chat`.
- **`AgentEvaluatorOperator`** joins predictions with ground-truth contexts to compute scores across all selected metrics.
- Built-in metrics in `packages/derisk-rag/src/derisk/rag/evaluation/` cover similarity, hit-rate, and relevancy, while [`packages/derisk-serve/src/derisk_serve/agent/evaluation/evaluation_metric.py`](https://github.com/derisk-ai/openderisk/blob/main/packages/derisk-serve/src/derisk_serve/agent/evaluation/evaluation_metric.py) provides intent and app-link verification.
- The modular design allows extension without modifying core pipeline code.

## Frequently Asked Questions

### How does the framework measure agent latency?

The `AgentOutputOperator` records execution time when calling `multi_agents.app_chat` and stores this value as `prediction_cost` within each `EvaluationResult`. This allows precise tracking of performance bottlenecks alongside accuracy metrics.

### Can I evaluate agents without writing custom code?

Yes. The `AgentEvaluator` class accepts standard datasets (list of dictionaries with `query` and `contexts` keys) and can run any registered metric by name. Pass pre-registered metrics like `RetrieverSimilarityMetric` to the `evaluate()` method without subclassing.

### What is the difference between `BaseEvaluationResult` and `EvaluationResult`?

`BaseEvaluationResult` (defined in [`packages/derisk-core/src/derisk/core/interface/evaluation.py`](https://github.com/derisk-ai/openderisk/blob/main/packages/derisk-core/src/derisk/core/interface/evaluation.py)) contains the raw metric output (score, passing flag, prediction text). `EvaluationResult` wraps this container with evaluation metadata including the metric name, original query, dataset entry, and latency cost.

### How do I add a metric that requires external API calls?

Subclass `EvaluationMetric` and implement `sync_compute`. While the method name implies synchronous execution, you can invoke asynchronous clients within the method body. The `AgentEvaluatorOperator` handles the async orchestration layer, calling your metric's compute method within the DAG's async context.