# How Reciprocal Rank Fusion Improves Search Results Over Single-Strategy BM25 in context-mode

> Discover how Reciprocal Rank Fusion boosts mksglu/context-mode search recall and ranking stability beyond single-strategy BM25. Learn about its powerful combination approach.

- Repository: [Mert Köseoğlu/context-mode](https://github.com/mksglu/context-mode)
- Tags: deep-dive
- Published: 2026-04-24

---

**Reciprocal Rank Fusion (RRF) combines BM25 and trigram-based search results using the formula `1/(60 + rank)` to deliver higher recall and more stable rankings than a single-strategy BM25 lookup.**

**context-mode** is a documentation search library that stores text in a SQLite FTS5 virtual table. While a standard BM25 lookup retrieves matches from a single Porter-stemmed tokenizer, Reciprocal Rank Fusion introduces a second retrieval strategy—trigram-based fuzzy matching—and merges both result sets into a unified ranking that surfaces relevant documents even when tokens differ or queries contain typos.

## The Limitations of Single-Strategy BM25

A plain **BM25** search in `context-mode` relies on SQLite’s FTS5 Porter tokenizer. This approach applies aggressive stemming and stop-word filtering, which can cause two significant problems:

- **Token mismatch**: Rare or misspelled words (e.g., “instagit” vs. “insta-git”) fail to match stemmed tokens in the index.
- **Rank instability**: A single noisy term can dominate the BM25 score, pushing highly relevant documents down the list or excluding them entirely.

The standard `search` method in [`src/store.ts`](https://github.com/mksglu/context-mode/blob/main/src/store.ts) executes only this one retrieval strategy, limiting results to documents that contain the exact stemmed tokens.

## How RRF Implements Dual-Strategy Retrieval

**Reciprocal Rank Fusion** solves these limitations by executing two independent queries and merging their results. According to the `context-mode` source code, the fusion layer operates on:

1. **Standard BM25**: Porter-stemmed token matches (`this.search`)
2. **Trigram fuzzy search**: Character n-gram matches that survive tokenization differences (`this.searchTrigram`)

The algorithm uses a constant **K = 60** (the standard RRF constant) and scores each document by summing `1 / (K + i + 1)` where `i` is the zero-based rank position in each respective list. This parameter-free approach requires no manual tuning—documents that rank highly in either list receive a significant score boost.

## The Fusion Algorithm Implementation

The core fusion logic resides in [`src/store.ts`](https://github.com/mksglu/context-mode/blob/main/src/store.ts) at **lines 949-990** within the private `#rrfSearch` method. The implementation follows four distinct steps:

1. **Dual execution**: It calls both `this.search` and `this.searchTrigram` with a fetch limit twice the user-requested `limit` (minimum 10).
2. **Score accumulation**: It creates a `Map` keyed by `source::title`. For each occurrence of a document in either result list, it adds the reciprocal rank contribution.
3. **Aggregation**: After processing both lists, the map values are sorted by accumulated score in descending order.
4. **Normalization**: The method returns results trimmed to the requested `limit`, with a synthetic `rank` property set to `-score` so that higher scores produce lower numeric ranks (consistent with ranking conventions).

```typescript
// Conceptual representation of the RRF scoring (K = 60)
const score = (bm25Rank !== -1 ? 1 / (60 + bm25Rank + 1) : 0) 
            + (trigramRank !== -1 ? 1 / (60 + trigramRank + 1) : 0);

```

This approach ensures that a document ranking first in the trigram list but twentieth in the BM25 list still receives a competitive fused score, smoothing out outliers and improving **ranking robustness**.

## From Fusion to Final Results: searchWithFallback

The public entry point **`searchWithFallback`** (lines **404-452** in [`src/store.ts`](https://github.com/mksglu/context-mode/blob/main/src/store.ts)) orchestrates the complete retrieval pipeline. It invokes `#rrfSearch` first, then applies a proximity-based reranking step to boost documents where query terms appear closer together. If RRF yields no hits, the system falls back to fuzzy-corrected queries.

Because `searchWithFallback` uses RRF as its primary retrieval path, every search operation benefits from the dual-strategy coverage without requiring explicit configuration from the caller.

## Practical Code Examples

### Executing an RRF-Powered Search

The typical public API automatically handles Reciprocal Rank Fusion when you use `searchWithFallback`:

```typescript
import { ContentStore } from "./src/store.js";

const store = await ContentStore.create();   // opens/creates the SQLite DB
await store.indexDirectory("docs");          // load markdown files

// Search for "insta-git configuration" – RRF runs automatically.
const results = store.searchWithFallback("insta-git configuration", 5);
for (const r of results) {
  console.log(`${r.title} (source: ${r.source})`);
}

```

*Behind the scenes*: `searchWithFallback` → `#rrfSearch` → `search` + `searchTrigram` → fusion (see [lines 949-990](https://github.com/mksglu/context-mode/blob/main/src/store.ts#L949-L990)).

### Comparing Single-Strategy vs. Fused Results

You can observe the recall difference by comparing BM25-only against the RRF-fused results:

```typescript
const bm25Only = store.search("context-mode CLI", 10); // single strategy
const rrf = store.searchWithFallback("context-mode CLI", 10); // fused

console.log("BM25 only:", bm25Only.map(r => r.title));
console.log("RRF fused:", rrf.map(r => r.title));

```

Typical output shows that `rrf` returns all BM25 hits plus additional relevant chunks caught only by the trigram index, such as misspelled or partial token matches.

## Summary

- **Reciprocal Rank Fusion** in `context-mode` merges BM25 (Porter-stemmed) and trigram-based search results to maximize recall.
- The fusion algorithm uses **K = 60** and the formula `1/(K + rank)` to combine rankings without requiring parameter tuning.
- Implementation lives in [`src/store.ts`](https://github.com/mksglu/context-mode/blob/main/src/store.ts) lines **949-990** (`#rrfSearch`) and is invoked by the public `searchWithFallback` method (lines **404-452**).
- RRF improves **coverage** for misspelled terms, increases **recall** by including documents from either index, and provides **ranking stability** against noisy BM25 scores.
- The approach remains efficient at **O(limit)** time complexity despite running two queries.

## Frequently Asked Questions

### What is the significance of the constant K = 60 in RRF?

The constant **K = 60** acts as a smoothing factor that prevents the reciprocal rank score from exploding when a document ranks near the top of a list. It ensures that differences in rank position matter more at the top of results than at the bottom, providing stable fusion across different query types without requiring dataset-specific tuning.

### How does Reciprocal Rank Fusion handle misspelled queries?

RRF handles misspellings by incorporating the **trigram-based fuzzy search** as a secondary retrieval strategy. While the BM25 index may miss a term like “instagit” because it survives neither stemming nor tokenization, the trigram index matches character sequences (e.g., “ins”, “nst”, “sta”) and retrieves the correct document. The fusion algorithm then boosts these results if they rank highly in the trigram list, even if absent from the BM25 results.

### What is the performance cost of RRF compared to BM25-only?

RRF executes two queries instead of one, but because both use SQLite FTS5 indices and the fusion step operates on a limited result set (fetching only `limit × 2` or 10 documents per query), the time complexity remains **O(limit)**. The additional overhead is negligible for typical documentation search scenarios, and the trade-off for improved recall and ranking stability is justified in the `context-mode` implementation.

### Should I use `searchWithFallback` or call `#rrfSearch` directly?

Always use **`searchWithFallback`**, the public API. The `#rrfSearch` method is private (denoted by the `#` prefix) and intended for internal use within the `ContentStore` class. `searchWithFallback` provides the complete pipeline including RRF fusion, proximity reranking, and fallback logic for zero-result queries, ensuring consistent behavior across the library.