How to Implement Custom Evaluation Metrics in PyLate: A Step-by-Step Guide

You can implement custom evaluation metrics in PyLate by creating a Ranx Metric subclass and passing it to the metrics argument of any PyLate evaluator, since PyLate delegates all metric computation to the underlying Ranx library.

PyLate is a neural search library that simplifies dense retrieval evaluation by integrating with the Ranx metrics library. When you need to track domain-specific performance indicators beyond standard NDCG or Recall, implementing custom evaluation metrics in PyLate allows you to extend the framework without modifying core evaluation logic.

How PyLate's Evaluation Pipeline Works

PyLate's evaluation utilities in pylate/evaluation/beir.py, pylate/evaluation/pylate_information_retrieval_evaluator.py, and pylate/evaluation/nano_beir_evaluator.py are thin wrappers around Ranx. The workflow follows five stages:

  1. Model inference generates ranked document IDs and scores for each query.
  2. Run conversion transforms these results into a ranx.Run object.
  3. Qrels conversion converts ground-truth relevance judgments into ranx.Qrels.
  4. Metric computation invokes ranx.evaluate with your specified metrics list.
  5. Results aggregation returns a dictionary of metric scores.

Because Ranx handles the actual mathematics, you only need to define a valid Ranx metric and pass it through PyLate's metrics parameter.

Creating Custom Evaluation Metrics in PyLate

Follow these steps to implement a custom metric, such as a specialized Precision@K variant with macro-averaging.

Step 1: Define a Custom Ranx Metric Class

Create a new file at pylate/evaluation/custom_metrics.py and subclass ranx.Metric. Implement the compute method to calculate your metric for a single query's relevance list.


# pylate/evaluation/custom_metrics.py

from typing import List
from ranx import Metric

class PrecisionAtK(Metric):
    """Precision@k with custom macro-averaging strategy."""
    
    def __init__(self, k: int):
        super().__init__(name=f"precision@{k}")
        self.k = k
    
    def compute(self, relevance: List[int]) -> float:
        # relevance is a binary list (1=relevant, 0=non-relevant) for top-k docs

        if len(relevance) == 0:
            return 0.0
        return sum(relevance[:self.k]) / self.k

Step 2: Export the Metric in PyLate's Evaluation Module

Make your metric importable by adding it to pylate/evaluation/__init__.py.


# pylate/evaluation/__init__.py

from .custom_metrics import PrecisionAtK

__all__ = [
    "load_custom_dataset",
    "PrecisionAtK",
    # ... other exports

]

Step 3: Instantiate and Use the Custom Metric

Pass your metric object to the metrics argument when calling any PyLate evaluator. You can mix custom objects with built-in Ranx string identifiers.

from pylate import evaluation
from pylate.evaluation import PrecisionAtK

# Prepare your data: scores, qrels, and queries

# scores = model.encode_and_retrieve(...)

# qrels = load_qrels(...)

# queries = load_queries(...)

# Create custom metric instance

custom_prec = PrecisionAtK(k=20)

# Evaluate with mixed metrics

results = evaluation.evaluate(
    scores=scores,
    qrels=qrels,
    queries=queries,
    metrics=[custom_prec, "ndcg@10", "hits@5", "map"],
)

print(f"Custom Precision@20: {results['precision@20']}")
print(f"NDCG@10: {results['ndcg@10']}")

Step 4: (Optional) Implement Custom Aggregation Logic

By default, Ranx averages per-query scores using the mean. To use a different aggregation strategy, such as weighted averaging, subclass ranx.Aggregation.

from ranx import Aggregation
from typing import List
import numpy as np

class WeightedMean(Aggregation):
    def __call__(self, per_query_scores: List[float]) -> float:
        weights = np.arange(1, len(per_query_scores) + 1)
        return np.average(per_query_scores, weights=weights)

class PrecisionAt20Weighted(PrecisionAtK):
    def __init__(self):
        super().__init__(k=20)
        self.aggregation = WeightedMean()

Key Files for Custom Metric Integration

Understanding where PyLate interfaces with Ranx helps you debug and extend the evaluation pipeline.

File Role in Custom Metric Integration
pylate/evaluation/beir.py Core evaluation function that forwards the metrics list to ranx.evaluate.
pylate/evaluation/pylate_information_retrieval_evaluator.py Sentence-transformers compatible evaluator that accepts custom metric objects.
pylate/evaluation/nano_beir_evaluator.py Lightweight BEIR evaluator demonstrating custom metric usage patterns.
pylate/evaluation/custom_metrics.py Recommended location for user-defined Metric subclasses.
pylate/evaluation/__init__.py Export custom metrics for convenient imports across the codebase.

Summary

Implementing custom evaluation metrics in PyLate requires minimal code changes because the library delegates metric computation to Ranx. The essential steps include:

  • Creating a ranx.Metric subclass in pylate/evaluation/custom_metrics.py with a compute method that processes binary relevance lists.
  • Exporting the metric in pylate/evaluation/__init__.py to enable clean imports.
  • Passing the metric instance to the metrics parameter in pylate.evaluation.evaluate alongside standard string identifiers like "ndcg@10".
  • Optionally overriding ranx.Aggregation to customize how per-query scores combine into final results.

Frequently Asked Questions

Can I use string identifiers for custom metrics in PyLate?

No, custom metrics must be passed as instantiated objects. Ranx recognizes built-in metrics like "ndcg@10" or "map" by string name, but user-defined subclasses of ranx.Metric must be instantiated and passed directly to the metrics list in pylate.evaluation.evaluate.

How does PyLate handle metric aggregation across queries?

By default, PyLate relies on Ranx's standard mean aggregation. Each metric's compute method returns a per-query score, and Ranx averages these using arithmetic mean unless you specify otherwise. To implement weighted or custom aggregation, subclass ranx.Aggregation and assign it to your metric's aggregation attribute.

Where should I store custom metrics in the PyLate codebase?

The recommended location is pylate/evaluation/custom_metrics.py. After defining your metric class there, expose it in pylate/evaluation/__init__.py by importing it and adding it to the __all__ list. This pattern keeps your extensions organized and accessible via from pylate.evaluation import YourMetric.

Can I mix custom metrics with built-in Ranx metrics?

Yes, the metrics parameter in PyLate's evaluation functions accepts a heterogeneous list containing both string identifiers for built-in metrics and instantiated custom metric objects. For example: metrics=[PrecisionAtK(k=20), "ndcg@10", "hits@5"] works correctly because Ranx processes each element according to its type.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →