How the OpenRisk Eval Framework Evaluates Agent Performance and Accuracy

The OpenRisk evaluation framework quantifies agent performance and accuracy by orchestrating a data-flow DAG that captures predictions via AgentOutputOperator, scores them against ground-truth contexts using pluggable EvaluationMetric implementations, and returns structured EvaluationResult objects containing numeric scores, pass/fail flags, and latency metrics.

The derisk-ai/openderisk repository provides a modular evaluation system designed to benchmark LLM-driven agents and retrievers. By combining abstract metric definitions with a directed acyclic graph (DAG) execution model, the framework enables precise measurement of both semantic relevance and task-specific correctness.

Core Evaluation Abstractions

The EvaluationMetric Contract

All accuracy measurements in OpenRisk derive from the abstract EvaluationMetric class defined in packages/derisk-core/src/derisk/core/interface/evaluation.py. This contract requires a single synchronous method:

def sync_compute(self, prediction, contexts, query=None) -> BaseEvaluationResult:

Concrete implementations override this method to define domain-specific scoring logic. For example, RetrieverSimilarityMetric computes cosine similarity between the agent's answer and ground-truth contexts, while IntentMetric parses JSON predictions to verify exact intent matches. Each metric returns a BaseEvaluationResult containing the raw score, a boolean passing flag, and optional feedback text.

Result Containers and Registry

The framework uses two primary data containers also located in packages/derisk-core/src/derisk/core/interface/evaluation.py:

  • BaseEvaluationResult: Holds the numeric score, prediction text, passing boolean, and contexts list for a single metric computation.
  • EvaluationResult: Wraps BaseEvaluationResult with additional metadata including the metric name, query string, raw dataset entry, and prediction_cost (latency).

The MetricManage class maintains a global registry mapping metric names to their classes. Developers register new metrics using metric_manage.register_metric(MyMetric), enabling dynamic lookup during evaluation runs.

The Agent Evaluation DAG

The evaluation pipeline is implemented as a data-flow DAG in packages/derisk-serve/src/derisk_serve/agent/evaluation/evaluation.py. Three specialized operators coordinate the workflow:

AgentOutputOperator (Prediction Capture)

AgentOutputOperator is a MapOperator that invokes the target agent via multi_agents.app_chat (defined in packages/derisk-serve/src/derisk_serve/agent/agents/controller.py). It captures the final answer and records execution time as prediction_cost. This operator transforms a raw query into an agent prediction ready for scoring.

AgentEvaluatorOperator (Scoring)

AgentEvaluatorOperator acts as a JoinOperator that receives the original query, the agent's prediction, and the ground-truth contexts. For each selected metric, it asynchronously executes metric.compute(prediction, contexts) and assembles the outputs into EvaluationResult objects.

AgentEvaluator (Orchestration)

The AgentEvaluator class (subclass of Evaluator) orchestrates the full pipeline:

  1. Dataset ingestion: An IteratorTrigger yields rows containing queries and contexts.
  2. Query extraction: A MapOperator extracts the query field specified by query_key.
  3. Agent invocation: AgentOutputOperator generates predictions.
  4. Metric evaluation: AgentEvaluatorOperator applies all registered metrics.
  5. Result aggregation: Returns a List[List[EvaluationResult]] where outer indices represent dataset rows and inner indices represent individual metrics.

Built-in Accuracy Metrics

OpenRisk ships with concrete metrics covering both relevance and correctness dimensions:

  • RetrieverSimilarityMetric: Uses embeddings to calculate cosine similarity between predictions and contexts.
  • RetrieverHitRateMetric: Binary score indicating whether the correct context appears in the top-k retrievals.
  • AnswerRelevancyMetric: Assesses semantic relevance of the generated answer to the query.
  • AppLinkMetric & IntentMetric: Parse JSON-encoded predictions to verify exact matches for app_name or intent fields against expected values.

These metrics are registered in packages/derisk-serve/src/derisk_serve/agent/evaluation/evaluation_metric.py and packages/derisk-rag/src/derisk/rag/evaluation/.

Running an Evaluation

To evaluate agent performance against a dataset with the default similarity metric:

import asyncio
from derisk.rag.evaluation import RetrieverSimilarityMetric
from derisk_serve.agent.evaluation.evaluation import AgentEvaluator, AgentOutputOperator

dataset = [
    {
        "query": "What does the OpenRisk API provide?",
        "contexts": [
            "OpenRisk offers risk-assessment APIs for financial models.",
        ],
    },
]

evaluator = AgentEvaluator(
    operator_cls=AgentOutputOperator,
    operator_kwargs={"app_code": "my_agent"},
)

results = asyncio.run(evaluator.evaluate(dataset))

# Inspect results: results[row_index][metric_index]

print(results[0][0].metric_name)   # "RetrieverSimilarityMetric"

print(results[0][0].score)         # Numeric similarity score

print(results[0][0].passing)       # True/False

print(results[0][0].prediction_cost)  # Latency in seconds

Extending the Framework

Add custom accuracy checks by subclassing EvaluationMetric and registering with MetricManage:

from derisk.core.interface.evaluation import EvaluationMetric, BaseEvaluationResult, metric_manage
import json

class IntentCorrectnessMetric(EvaluationMetric[str, str]):
    def sync_compute(self, prediction, contexts, query=None):
        try:
            parsed = json.loads(prediction)
            intent = parsed.get("intent")
            score = 1.0 if intent in contexts else 0.0
            passing = bool(score)
        except Exception:
            score = 0.0
            passing = False
            intent = None
        
        return BaseEvaluationResult(
            score=score, 
            prediction=intent, 
            passing=passing
        )

# Register for dynamic discovery

metric_manage.register_metric(IntentCorrectnessMetric)

# Use in evaluation

custom_evaluator = AgentEvaluator(
    operator_cls=AgentOutputOperator,
    operator_kwargs={"app_code": "my_agent"},
)
results = asyncio.run(custom_evaluator.evaluate(
    dataset, 
    metrics=[IntentCorrectnessMetric()]
))

Summary

  • The OpenRisk evaluation framework uses a DAG architecture to assess agent performance and accuracy through composable operators.
  • EvaluationMetric defines the scoring contract via sync_compute, while MetricManage enables dynamic metric registration.
  • AgentOutputOperator captures predictions and latency (prediction_cost) by calling multi_agents.app_chat.
  • AgentEvaluatorOperator joins predictions with ground-truth contexts to compute scores across all selected metrics.
  • Built-in metrics in packages/derisk-rag/src/derisk/rag/evaluation/ cover similarity, hit-rate, and relevancy, while packages/derisk-serve/src/derisk_serve/agent/evaluation/evaluation_metric.py provides intent and app-link verification.
  • The modular design allows extension without modifying core pipeline code.

Frequently Asked Questions

How does the framework measure agent latency?

The AgentOutputOperator records execution time when calling multi_agents.app_chat and stores this value as prediction_cost within each EvaluationResult. This allows precise tracking of performance bottlenecks alongside accuracy metrics.

Can I evaluate agents without writing custom code?

Yes. The AgentEvaluator class accepts standard datasets (list of dictionaries with query and contexts keys) and can run any registered metric by name. Pass pre-registered metrics like RetrieverSimilarityMetric to the evaluate() method without subclassing.

What is the difference between BaseEvaluationResult and EvaluationResult?

BaseEvaluationResult (defined in packages/derisk-core/src/derisk/core/interface/evaluation.py) contains the raw metric output (score, passing flag, prediction text). EvaluationResult wraps this container with evaluation metadata including the metric name, original query, dataset entry, and latency cost.

How do I add a metric that requires external API calls?

Subclass EvaluationMetric and implement sync_compute. While the method name implies synchronous execution, you can invoke asynchronous clients within the method body. The AgentEvaluatorOperator handles the async orchestration layer, calling your metric's compute method within the DAG's async context.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →