How the OpenRisk Eval Framework Evaluates Agent Performance and Accuracy
The OpenRisk evaluation framework quantifies agent performance and accuracy by orchestrating a data-flow DAG that captures predictions via AgentOutputOperator, scores them against ground-truth contexts using pluggable EvaluationMetric implementations, and returns structured EvaluationResult objects containing numeric scores, pass/fail flags, and latency metrics.
The derisk-ai/openderisk repository provides a modular evaluation system designed to benchmark LLM-driven agents and retrievers. By combining abstract metric definitions with a directed acyclic graph (DAG) execution model, the framework enables precise measurement of both semantic relevance and task-specific correctness.
Core Evaluation Abstractions
The EvaluationMetric Contract
All accuracy measurements in OpenRisk derive from the abstract EvaluationMetric class defined in packages/derisk-core/src/derisk/core/interface/evaluation.py. This contract requires a single synchronous method:
def sync_compute(self, prediction, contexts, query=None) -> BaseEvaluationResult:
Concrete implementations override this method to define domain-specific scoring logic. For example, RetrieverSimilarityMetric computes cosine similarity between the agent's answer and ground-truth contexts, while IntentMetric parses JSON predictions to verify exact intent matches. Each metric returns a BaseEvaluationResult containing the raw score, a boolean passing flag, and optional feedback text.
Result Containers and Registry
The framework uses two primary data containers also located in packages/derisk-core/src/derisk/core/interface/evaluation.py:
BaseEvaluationResult: Holds the numericscore,predictiontext,passingboolean, andcontextslist for a single metric computation.EvaluationResult: WrapsBaseEvaluationResultwith additional metadata including the metric name, query string, raw dataset entry, andprediction_cost(latency).
The MetricManage class maintains a global registry mapping metric names to their classes. Developers register new metrics using metric_manage.register_metric(MyMetric), enabling dynamic lookup during evaluation runs.
The Agent Evaluation DAG
The evaluation pipeline is implemented as a data-flow DAG in packages/derisk-serve/src/derisk_serve/agent/evaluation/evaluation.py. Three specialized operators coordinate the workflow:
AgentOutputOperator (Prediction Capture)
AgentOutputOperator is a MapOperator that invokes the target agent via multi_agents.app_chat (defined in packages/derisk-serve/src/derisk_serve/agent/agents/controller.py). It captures the final answer and records execution time as prediction_cost. This operator transforms a raw query into an agent prediction ready for scoring.
AgentEvaluatorOperator (Scoring)
AgentEvaluatorOperator acts as a JoinOperator that receives the original query, the agent's prediction, and the ground-truth contexts. For each selected metric, it asynchronously executes metric.compute(prediction, contexts) and assembles the outputs into EvaluationResult objects.
AgentEvaluator (Orchestration)
The AgentEvaluator class (subclass of Evaluator) orchestrates the full pipeline:
- Dataset ingestion: An
IteratorTriggeryields rows containing queries and contexts. - Query extraction: A
MapOperatorextracts the query field specified byquery_key. - Agent invocation:
AgentOutputOperatorgenerates predictions. - Metric evaluation:
AgentEvaluatorOperatorapplies all registered metrics. - Result aggregation: Returns a
List[List[EvaluationResult]]where outer indices represent dataset rows and inner indices represent individual metrics.
Built-in Accuracy Metrics
OpenRisk ships with concrete metrics covering both relevance and correctness dimensions:
- RetrieverSimilarityMetric: Uses embeddings to calculate cosine similarity between predictions and contexts.
- RetrieverHitRateMetric: Binary score indicating whether the correct context appears in the top-k retrievals.
- AnswerRelevancyMetric: Assesses semantic relevance of the generated answer to the query.
- AppLinkMetric & IntentMetric: Parse JSON-encoded predictions to verify exact matches for
app_nameorintentfields against expected values.
These metrics are registered in packages/derisk-serve/src/derisk_serve/agent/evaluation/evaluation_metric.py and packages/derisk-rag/src/derisk/rag/evaluation/.
Running an Evaluation
To evaluate agent performance against a dataset with the default similarity metric:
import asyncio
from derisk.rag.evaluation import RetrieverSimilarityMetric
from derisk_serve.agent.evaluation.evaluation import AgentEvaluator, AgentOutputOperator
dataset = [
{
"query": "What does the OpenRisk API provide?",
"contexts": [
"OpenRisk offers risk-assessment APIs for financial models.",
],
},
]
evaluator = AgentEvaluator(
operator_cls=AgentOutputOperator,
operator_kwargs={"app_code": "my_agent"},
)
results = asyncio.run(evaluator.evaluate(dataset))
# Inspect results: results[row_index][metric_index]
print(results[0][0].metric_name) # "RetrieverSimilarityMetric"
print(results[0][0].score) # Numeric similarity score
print(results[0][0].passing) # True/False
print(results[0][0].prediction_cost) # Latency in seconds
Extending the Framework
Add custom accuracy checks by subclassing EvaluationMetric and registering with MetricManage:
from derisk.core.interface.evaluation import EvaluationMetric, BaseEvaluationResult, metric_manage
import json
class IntentCorrectnessMetric(EvaluationMetric[str, str]):
def sync_compute(self, prediction, contexts, query=None):
try:
parsed = json.loads(prediction)
intent = parsed.get("intent")
score = 1.0 if intent in contexts else 0.0
passing = bool(score)
except Exception:
score = 0.0
passing = False
intent = None
return BaseEvaluationResult(
score=score,
prediction=intent,
passing=passing
)
# Register for dynamic discovery
metric_manage.register_metric(IntentCorrectnessMetric)
# Use in evaluation
custom_evaluator = AgentEvaluator(
operator_cls=AgentOutputOperator,
operator_kwargs={"app_code": "my_agent"},
)
results = asyncio.run(custom_evaluator.evaluate(
dataset,
metrics=[IntentCorrectnessMetric()]
))
Summary
- The OpenRisk evaluation framework uses a DAG architecture to assess agent performance and accuracy through composable operators.
EvaluationMetricdefines the scoring contract viasync_compute, whileMetricManageenables dynamic metric registration.AgentOutputOperatorcaptures predictions and latency (prediction_cost) by callingmulti_agents.app_chat.AgentEvaluatorOperatorjoins predictions with ground-truth contexts to compute scores across all selected metrics.- Built-in metrics in
packages/derisk-rag/src/derisk/rag/evaluation/cover similarity, hit-rate, and relevancy, whilepackages/derisk-serve/src/derisk_serve/agent/evaluation/evaluation_metric.pyprovides intent and app-link verification. - The modular design allows extension without modifying core pipeline code.
Frequently Asked Questions
How does the framework measure agent latency?
The AgentOutputOperator records execution time when calling multi_agents.app_chat and stores this value as prediction_cost within each EvaluationResult. This allows precise tracking of performance bottlenecks alongside accuracy metrics.
Can I evaluate agents without writing custom code?
Yes. The AgentEvaluator class accepts standard datasets (list of dictionaries with query and contexts keys) and can run any registered metric by name. Pass pre-registered metrics like RetrieverSimilarityMetric to the evaluate() method without subclassing.
What is the difference between BaseEvaluationResult and EvaluationResult?
BaseEvaluationResult (defined in packages/derisk-core/src/derisk/core/interface/evaluation.py) contains the raw metric output (score, passing flag, prediction text). EvaluationResult wraps this container with evaluation metadata including the metric name, original query, dataset entry, and latency cost.
How do I add a metric that requires external API calls?
Subclass EvaluationMetric and implement sync_compute. While the method name implies synchronous execution, you can invoke asynchronous clients within the method body. The AgentEvaluatorOperator handles the async orchestration layer, calling your metric's compute method within the DAG's async context.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →