# How to Implement Cross-Encoder Reranking for Improved Retrieval Precision in OpenRisk

> Implement cross-encoder reranking in OpenRisk to reorder retrieved text chunks by predicted relevance. Improve retrieval precision before LLM generation with Sentence-Transformers.

- Repository: [derisk-ai/openderisk](https://github.com/derisk-ai/openderisk)
- Tags: how-to-guide
- Published: 2026-02-28

---

**Use the `CrossEncoderRanker` class from [`packages/derisk-core/src/derisk/rag/retriever/rerank.py`](https://github.com/derisk-ai/openderisk/blob/main/packages/derisk-core/src/derisk/rag/retriever/rerank.py) to rescore retrieved chunks using a Sentence-Transformers cross-encoder model, reorder candidates by predicted relevance, and return the top-k results before LLM generation.**

The OpenRisk repository provides a modular RAG (Retrieval-Augmented Generation) pipeline where retrieved chunks undergo optional reranking to boost precision. Implementing **cross-encoder reranking** allows you to score query-document pairs directly, yielding superior relevance compared to initial vector or keyword search alone.

## Architecture and Implementation Details

The reranking system in OpenRisk follows a clean abstraction pattern that enables swapping ranking strategies without touching downstream code.

### The Ranker Base Class

At `lines 14-73` of [`packages/derisk-core/src/derisk/rag/retriever/rerank.py`](https://github.com/derisk-ai/openderisk/blob/main/packages/derisk-core/src/derisk/rag/retriever/rerank.py), the abstract `Ranker` class defines the common interface for all rerankers. It specifies the `rank` and `arank` methods that concrete implementations must override, and supplies helper utilities including `_filter` and `_rerank_with_scores`. The async support at `lines 43-60` uses `blocking_func_to_async_no_executor` to run synchronous ranking logic in a background thread, ensuring safe integration with async workflows.

### CrossEncoderRanker Implementation

The `CrossEncoderRanker` class (starting at `line 77`) wraps a Sentence-Transformers **CrossEncoder** model. Its constructor (`lines 86-111`) instantiates the model with a configurable device (CPU or CUDA) and a fixed `max_length` of 512 tokens. The class is registered as a first-class resource via the `@register_resource` decorator (`lines 174-182`), making it selectable through the UI or YAML configuration without code changes.

## How Cross-Encoder Reranking Works

The reranking process involves three distinct phases executed within the `rank` method (`lines 122-160`):

1. **Model Loading** – The constructor imports `CrossEncoder` from *sentence-transformers* and initializes it with your specified model name and device:

   ```python
   self._model = CrossEncoder(model, max_length=512, device=device)
   ```

   If the dependency is missing, an explicit `ImportError` is raised at `lines 124-127` to guide installation.

2. **Scoring Query-Document Pairs** – The `rank` method constructs `[query, document]` pairs for every candidate chunk and passes them to the model:

   ```python
   query_content_pairs = [[query or "", content or ""] for content in contents]
   rank_scores = self._model.predict(sentences=query_content_pairs)
   ```

   This occurs at `lines 140-152`, where the cross-encoder outputs a scalar relevance score for each pair.

3. **Reordering and Truncation** – Each candidate's `score` attribute is overwritten with the predicted relevance, the list is sorted in descending order, and truncated to `topk` (`lines 154-160`).

## Configuration and Usage

Because `CrossEncoderRanker` inherits from `Ranker`, it integrates seamlessly with existing retrieval pipelines. You can invoke it synchronously for scripts or asynchronously for web applications.

### Synchronous Implementation

Use this pattern in batch processing scripts or Jupyter notebooks:

```python
from derisk.rag.retriever.rerank import CrossEncoderRanker
from derisk.core import Chunk

# chunks is a list[Chunk] from your initial retriever (BM25 or dense)

chunks = [...]
query = "What are the regulatory requirements for crypto exchanges?"

ranker = CrossEncoderRanker(
    topk=5,
    model="BAAI/bge-reranker-base",
    device="cpu"
)

top_chunks = ranker.rank(candidates_with_scores=chunks, query=query)

```

### Asynchronous Implementation

For FastAPI or other async frameworks, use `arank` to avoid blocking the event loop:

```python
import asyncio
from derisk.rag.retriever.rerank import CrossEncoderRanker

async def rerank_documents(chunks: list, query: str):
    ranker = CrossEncoderRanker(
        topk=5, 
        model="BAAI/bge-reranker-base", 
        device="cpu"
    )
    # Runs sync rank() in a thread pool

    return await ranker.arank(candidates_with_scores=chunks, query=query)

# Usage

# results = asyncio.run(rerank_documents(chunks, query))

```

### YAML Configuration

Since the ranker is registered as a resource, you can enable it via configuration files without modifying code:

```yaml
ranker:
  type: cross_encoder_ranker
  topk: 5
  model: BAAI/bge-reranker-base
  device: cpu

```

The system instantiates `CrossEncoderRanker` automatically using these parameters when the pipeline initializes.

## Performance Considerations and Trade-offs

**Cross-encoder reranking** delivers higher precision than score-based rankers like `DefaultRanker` or `RRFRanker`, but introduces computational overhead. Consider these optimizations:

- **Model Selection**: Use lightweight models such as `BAAI/bge-reranker-base` rather than large generative models to minimize latency.
- **Top-k Filtering**: Set `topk` to a small number (e.g., 5) to limit the number of expensive forward passes.
- **Device Placement**: Specify `device="cuda"` when GPU resources are available to accelerate inference.

## Summary

- The `CrossEncoderRanker` class in [`packages/derisk-core/src/derisk/rag/retriever/rerank.py`](https://github.com/derisk-ai/openderisk/blob/main/packages/derisk-core/src/derisk/rag/retriever/rerank.py) implements high-precision reranking by scoring query-document pairs with a Sentence-Transformers model.
- It inherits from the `Ranker` ABC, providing `rank` (synchronous) and `arank` (asynchronous) methods for flexible integration.
- Configuration occurs via the `@register_resource` decorator, supporting both code-based and YAML-based setup.
- The implementation requires `sentence-transformers` and handles missing dependencies with explicit `ImportError` messages.
- For production use, select lightweight models and small `topk` values to balance relevance gains against latency costs.

## Frequently Asked Questions

### What is the difference between a cross-encoder and a bi-encoder?

A **bi-encoder** embeds queries and documents separately into a shared vector space, allowing fast similarity computation via dot product or cosine similarity. A **cross-encoder** concatenates the query and document as input to a transformer model, computing a direct relevance score through self-attention across both texts. Cross-encoders achieve higher accuracy but are slower because they require a forward pass for every query-document pair, making them suitable for reranking a small candidate set rather than indexing an entire corpus.

### Which model should I use for cross-encoder reranking in OpenRisk?

The OpenRisk implementation accepts any Sentence-Transformers `CrossEncoder` model. For general-purpose reranking, `BAAI/bge-reranker-base` provides an optimal balance between accuracy and inference speed. If you have domain-specific data (e.g., financial or legal text), fine-tune a cross-encoder on your corpus and pass the custom model name to the `model` parameter in `CrossEncoderRanker.__init__`.

### How do I handle asynchronous reranking in FastAPI endpoints?

Import `CrossEncoderRanker` and call the `arank` method instead of `rank`. According to the source code at `lines 43-60` of [`rerank.py`](https://github.com/derisk-ai/openderisk/blob/main/rerank.py), `arank` wraps the synchronous `rank` method using `blocking_func_to_async_no_executor`, which executes the model inference in a background thread pool. This prevents the CPU-intensive reranking from blocking your event loop while remaining compatible with async database and HTTP operations.

### What happens if sentence-transformers is not installed?

If the `sentence-transformers` package is missing when `CrossEncoderRanker` initializes, the code raises a clear `ImportError` at `lines 124-127` with instructions to install the dependency. The repository declares this as an optional dependency in its packaging metadata, so you must explicitly install it using `pip install sentence-transformers` before using the cross-encoder functionality.