How to Implement Custom Evaluation Metrics in PyLate: A Step-by-Step Guide
You can implement custom evaluation metrics in PyLate by creating a Ranx Metric subclass and passing it to the metrics argument of any PyLate evaluator, since PyLate delegates all metric computation to the underlying Ranx library.
PyLate is a neural search library that simplifies dense retrieval evaluation by integrating with the Ranx metrics library. When you need to track domain-specific performance indicators beyond standard NDCG or Recall, implementing custom evaluation metrics in PyLate allows you to extend the framework without modifying core evaluation logic.
How PyLate's Evaluation Pipeline Works
PyLate's evaluation utilities in pylate/evaluation/beir.py, pylate/evaluation/pylate_information_retrieval_evaluator.py, and pylate/evaluation/nano_beir_evaluator.py are thin wrappers around Ranx. The workflow follows five stages:
- Model inference generates ranked document IDs and scores for each query.
- Run conversion transforms these results into a
ranx.Runobject. - Qrels conversion converts ground-truth relevance judgments into
ranx.Qrels. - Metric computation invokes
ranx.evaluatewith your specified metrics list. - Results aggregation returns a dictionary of metric scores.
Because Ranx handles the actual mathematics, you only need to define a valid Ranx metric and pass it through PyLate's metrics parameter.
Creating Custom Evaluation Metrics in PyLate
Follow these steps to implement a custom metric, such as a specialized Precision@K variant with macro-averaging.
Step 1: Define a Custom Ranx Metric Class
Create a new file at pylate/evaluation/custom_metrics.py and subclass ranx.Metric. Implement the compute method to calculate your metric for a single query's relevance list.
# pylate/evaluation/custom_metrics.py
from typing import List
from ranx import Metric
class PrecisionAtK(Metric):
"""Precision@k with custom macro-averaging strategy."""
def __init__(self, k: int):
super().__init__(name=f"precision@{k}")
self.k = k
def compute(self, relevance: List[int]) -> float:
# relevance is a binary list (1=relevant, 0=non-relevant) for top-k docs
if len(relevance) == 0:
return 0.0
return sum(relevance[:self.k]) / self.k
Step 2: Export the Metric in PyLate's Evaluation Module
Make your metric importable by adding it to pylate/evaluation/__init__.py.
# pylate/evaluation/__init__.py
from .custom_metrics import PrecisionAtK
__all__ = [
"load_custom_dataset",
"PrecisionAtK",
# ... other exports
]
Step 3: Instantiate and Use the Custom Metric
Pass your metric object to the metrics argument when calling any PyLate evaluator. You can mix custom objects with built-in Ranx string identifiers.
from pylate import evaluation
from pylate.evaluation import PrecisionAtK
# Prepare your data: scores, qrels, and queries
# scores = model.encode_and_retrieve(...)
# qrels = load_qrels(...)
# queries = load_queries(...)
# Create custom metric instance
custom_prec = PrecisionAtK(k=20)
# Evaluate with mixed metrics
results = evaluation.evaluate(
scores=scores,
qrels=qrels,
queries=queries,
metrics=[custom_prec, "ndcg@10", "hits@5", "map"],
)
print(f"Custom Precision@20: {results['precision@20']}")
print(f"NDCG@10: {results['ndcg@10']}")
Step 4: (Optional) Implement Custom Aggregation Logic
By default, Ranx averages per-query scores using the mean. To use a different aggregation strategy, such as weighted averaging, subclass ranx.Aggregation.
from ranx import Aggregation
from typing import List
import numpy as np
class WeightedMean(Aggregation):
def __call__(self, per_query_scores: List[float]) -> float:
weights = np.arange(1, len(per_query_scores) + 1)
return np.average(per_query_scores, weights=weights)
class PrecisionAt20Weighted(PrecisionAtK):
def __init__(self):
super().__init__(k=20)
self.aggregation = WeightedMean()
Key Files for Custom Metric Integration
Understanding where PyLate interfaces with Ranx helps you debug and extend the evaluation pipeline.
| File | Role in Custom Metric Integration |
|---|---|
pylate/evaluation/beir.py |
Core evaluation function that forwards the metrics list to ranx.evaluate. |
pylate/evaluation/pylate_information_retrieval_evaluator.py |
Sentence-transformers compatible evaluator that accepts custom metric objects. |
pylate/evaluation/nano_beir_evaluator.py |
Lightweight BEIR evaluator demonstrating custom metric usage patterns. |
pylate/evaluation/custom_metrics.py |
Recommended location for user-defined Metric subclasses. |
pylate/evaluation/__init__.py |
Export custom metrics for convenient imports across the codebase. |
Summary
Implementing custom evaluation metrics in PyLate requires minimal code changes because the library delegates metric computation to Ranx. The essential steps include:
- Creating a
ranx.Metricsubclass inpylate/evaluation/custom_metrics.pywith acomputemethod that processes binary relevance lists. - Exporting the metric in
pylate/evaluation/__init__.pyto enable clean imports. - Passing the metric instance to the
metricsparameter inpylate.evaluation.evaluatealongside standard string identifiers like"ndcg@10". - Optionally overriding
ranx.Aggregationto customize how per-query scores combine into final results.
Frequently Asked Questions
Can I use string identifiers for custom metrics in PyLate?
No, custom metrics must be passed as instantiated objects. Ranx recognizes built-in metrics like "ndcg@10" or "map" by string name, but user-defined subclasses of ranx.Metric must be instantiated and passed directly to the metrics list in pylate.evaluation.evaluate.
How does PyLate handle metric aggregation across queries?
By default, PyLate relies on Ranx's standard mean aggregation. Each metric's compute method returns a per-query score, and Ranx averages these using arithmetic mean unless you specify otherwise. To implement weighted or custom aggregation, subclass ranx.Aggregation and assign it to your metric's aggregation attribute.
Where should I store custom metrics in the PyLate codebase?
The recommended location is pylate/evaluation/custom_metrics.py. After defining your metric class there, expose it in pylate/evaluation/__init__.py by importing it and adding it to the __all__ list. This pattern keeps your extensions organized and accessible via from pylate.evaluation import YourMetric.
Can I mix custom metrics with built-in Ranx metrics?
Yes, the metrics parameter in PyLate's evaluation functions accepts a heterogeneous list containing both string identifiers for built-in metrics and instantiated custom metric objects. For example: metrics=[PrecisionAtK(k=20), "ndcg@10", "hits@5"] works correctly because Ranx processes each element according to its type.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →