# How to Implement Custom Evaluation Metrics in PyLate: A Step-by-Step Guide

> Implement custom evaluation metrics in PyLate by creating a Ranx Metric subclass. Learn the steps to define and pass your own metrics to PyLate evaluators for precise performance tracking.

- Repository: [LightOn/pylate](https://github.com/lightonai/pylate)
- Tags: how-to-guide
- Published: 2026-03-06

---

**You can implement custom evaluation metrics in PyLate by creating a Ranx Metric subclass and passing it to the `metrics` argument of any PyLate evaluator, since PyLate delegates all metric computation to the underlying Ranx library.**

PyLate is a neural search library that simplifies dense retrieval evaluation by integrating with the **Ranx** metrics library. When you need to track domain-specific performance indicators beyond standard NDCG or Recall, implementing custom evaluation metrics in PyLate allows you to extend the framework without modifying core evaluation logic.

## How PyLate's Evaluation Pipeline Works

PyLate's evaluation utilities in [`pylate/evaluation/beir.py`](https://github.com/lightonai/pylate/blob/main/pylate/evaluation/beir.py), [`pylate/evaluation/pylate_information_retrieval_evaluator.py`](https://github.com/lightonai/pylate/blob/main/pylate/evaluation/pylate_information_retrieval_evaluator.py), and [`pylate/evaluation/nano_beir_evaluator.py`](https://github.com/lightonai/pylate/blob/main/pylate/evaluation/nano_beir_evaluator.py) are thin wrappers around Ranx. The workflow follows five stages:

1. **Model inference** generates ranked document IDs and scores for each query.
2. **Run conversion** transforms these results into a `ranx.Run` object.
3. **Qrels conversion** converts ground-truth relevance judgments into `ranx.Qrels`.
4. **Metric computation** invokes `ranx.evaluate` with your specified metrics list.
5. **Results aggregation** returns a dictionary of metric scores.

Because Ranx handles the actual mathematics, you only need to define a valid Ranx metric and pass it through PyLate's `metrics` parameter.

## Creating Custom Evaluation Metrics in PyLate

Follow these steps to implement a custom metric, such as a specialized Precision@K variant with macro-averaging.

### Step 1: Define a Custom Ranx Metric Class

Create a new file at [`pylate/evaluation/custom_metrics.py`](https://github.com/lightonai/pylate/blob/main/pylate/evaluation/custom_metrics.py) and subclass `ranx.Metric`. Implement the `compute` method to calculate your metric for a single query's relevance list.

```python

# pylate/evaluation/custom_metrics.py

from typing import List
from ranx import Metric

class PrecisionAtK(Metric):
    """Precision@k with custom macro-averaging strategy."""
    
    def __init__(self, k: int):
        super().__init__(name=f"precision@{k}")
        self.k = k
    
    def compute(self, relevance: List[int]) -> float:
        # relevance is a binary list (1=relevant, 0=non-relevant) for top-k docs

        if len(relevance) == 0:
            return 0.0
        return sum(relevance[:self.k]) / self.k

```

### Step 2: Export the Metric in PyLate's Evaluation Module

Make your metric importable by adding it to [`pylate/evaluation/__init__.py`](https://github.com/lightonai/pylate/blob/main/pylate/evaluation/__init__.py).

```python

# pylate/evaluation/__init__.py

from .custom_metrics import PrecisionAtK

__all__ = [
    "load_custom_dataset",
    "PrecisionAtK",
    # ... other exports

]

```

### Step 3: Instantiate and Use the Custom Metric

Pass your metric object to the `metrics` argument when calling any PyLate evaluator. You can mix custom objects with built-in Ranx string identifiers.

```python
from pylate import evaluation
from pylate.evaluation import PrecisionAtK

# Prepare your data: scores, qrels, and queries

# scores = model.encode_and_retrieve(...)

# qrels = load_qrels(...)

# queries = load_queries(...)

# Create custom metric instance

custom_prec = PrecisionAtK(k=20)

# Evaluate with mixed metrics

results = evaluation.evaluate(
    scores=scores,
    qrels=qrels,
    queries=queries,
    metrics=[custom_prec, "ndcg@10", "hits@5", "map"],
)

print(f"Custom Precision@20: {results['precision@20']}")
print(f"NDCG@10: {results['ndcg@10']}")

```

### Step 4: (Optional) Implement Custom Aggregation Logic

By default, Ranx averages per-query scores using the mean. To use a different aggregation strategy, such as weighted averaging, subclass `ranx.Aggregation`.

```python
from ranx import Aggregation
from typing import List
import numpy as np

class WeightedMean(Aggregation):
    def __call__(self, per_query_scores: List[float]) -> float:
        weights = np.arange(1, len(per_query_scores) + 1)
        return np.average(per_query_scores, weights=weights)

class PrecisionAt20Weighted(PrecisionAtK):
    def __init__(self):
        super().__init__(k=20)
        self.aggregation = WeightedMean()

```

## Key Files for Custom Metric Integration

Understanding where PyLate interfaces with Ranx helps you debug and extend the evaluation pipeline.

| File | Role in Custom Metric Integration |
|------|-----------------------------------|
| [`pylate/evaluation/beir.py`](https://github.com/lightonai/pylate/blob/main/pylate/evaluation/beir.py) | Core evaluation function that forwards the `metrics` list to `ranx.evaluate`. |
| [`pylate/evaluation/pylate_information_retrieval_evaluator.py`](https://github.com/lightonai/pylate/blob/main/pylate/evaluation/pylate_information_retrieval_evaluator.py) | Sentence-transformers compatible evaluator that accepts custom metric objects. |
| [`pylate/evaluation/nano_beir_evaluator.py`](https://github.com/lightonai/pylate/blob/main/pylate/evaluation/nano_beir_evaluator.py) | Lightweight BEIR evaluator demonstrating custom metric usage patterns. |
| [`pylate/evaluation/custom_metrics.py`](https://github.com/lightonai/pylate/blob/main/pylate/evaluation/custom_metrics.py) | Recommended location for user-defined `Metric` subclasses. |
| [`pylate/evaluation/__init__.py`](https://github.com/lightonai/pylate/blob/main/pylate/evaluation/__init__.py) | Export custom metrics for convenient imports across the codebase. |

## Summary

Implementing custom evaluation metrics in PyLate requires minimal code changes because the library delegates metric computation to Ranx. The essential steps include:

- Creating a `ranx.Metric` subclass in [`pylate/evaluation/custom_metrics.py`](https://github.com/lightonai/pylate/blob/main/pylate/evaluation/custom_metrics.py) with a `compute` method that processes binary relevance lists.
- Exporting the metric in [`pylate/evaluation/__init__.py`](https://github.com/lightonai/pylate/blob/main/pylate/evaluation/__init__.py) to enable clean imports.
- Passing the metric instance to the `metrics` parameter in `pylate.evaluation.evaluate` alongside standard string identifiers like `"ndcg@10"`.
- Optionally overriding `ranx.Aggregation` to customize how per-query scores combine into final results.

## Frequently Asked Questions

### Can I use string identifiers for custom metrics in PyLate?

No, custom metrics must be passed as instantiated objects. Ranx recognizes built-in metrics like `"ndcg@10"` or `"map"` by string name, but user-defined subclasses of `ranx.Metric` must be instantiated and passed directly to the `metrics` list in `pylate.evaluation.evaluate`.

### How does PyLate handle metric aggregation across queries?

By default, PyLate relies on Ranx's standard mean aggregation. Each metric's `compute` method returns a per-query score, and Ranx averages these using arithmetic mean unless you specify otherwise. To implement weighted or custom aggregation, subclass `ranx.Aggregation` and assign it to your metric's `aggregation` attribute.

### Where should I store custom metrics in the PyLate codebase?

The recommended location is [`pylate/evaluation/custom_metrics.py`](https://github.com/lightonai/pylate/blob/main/pylate/evaluation/custom_metrics.py). After defining your metric class there, expose it in [`pylate/evaluation/__init__.py`](https://github.com/lightonai/pylate/blob/main/pylate/evaluation/__init__.py) by importing it and adding it to the `__all__` list. This pattern keeps your extensions organized and accessible via `from pylate.evaluation import YourMetric`.

### Can I mix custom metrics with built-in Ranx metrics?

Yes, the `metrics` parameter in PyLate's evaluation functions accepts a heterogeneous list containing both string identifiers for built-in metrics and instantiated custom metric objects. For example: `metrics=[PrecisionAtK(k=20), "ndcg@10", "hits@5"]` works correctly because Ranx processes each element according to its type.