# How QMD Uses Position-Aware Blending for Intelligent Reranking

> Discover how QMD employs position-aware blending to intelligently rerank search results by fusing retrieval and neural scores. Learn about its tiered weighting system.

- Repository: [Tobias Lütke/qmd](https://github.com/tobi/qmd)
- Tags: internals
- Published: 2026-02-16

---

**QMD blends retrieval scores with neural reranker scores using a tiered weighting system based on Reciprocal Rank Fusion (RRF) position, preserving high-confidence exact matches while allowing semantic reordering.**

The `tobi/qmd` repository implements a sophisticated hybrid search pipeline that combines sparse retrieval (BM25), dense vector search, and cross-encoder reranking. Central to this architecture is a **position-aware blending** technique that prevents high-quality lexical matches from being buried by neural reranker scores, particularly when query expansion introduces semantic drift.

## The Hybrid Retrieval Pipeline

QMD begins by retrieving candidate documents using a hybrid approach that fuses BM25 and vector search results. These candidates are ordered by **Reciprocal Rank Fusion (RRF)**, which computes a single *RRF score* from the individual retrieval rankings and assigns each document an **RRF rank** where 1 represents the highest-confidence match.

A cross-encoder reranker (such as Qwen3-Reranker) then evaluates these same candidates, producing a **reranker score** normalized between 0 and 1. The challenge arises in merging these two signals without allowing the neural model to override strong lexical evidence, especially when the expanded query terms do not match the original search terms.

## Position-Aware Blending Strategy

### Tiered Weighting by RRF Rank

QMD implements position-aware blending by applying different **RRF weights** based on the original RRF rank tier. Higher-ranked documents retain more influence from the retrieval score, ensuring that top lexical matches are preserved:

| RRF rank | RRF weight |
|----------|------------|
| 1 – 3    | 0.75       |
| 4 – 10   | 0.60       |
| 11+      | 0.40       |

This tiered approach acknowledges that documents ranked highly by RRF are likely strong matches that should not be easily displaced by the reranker's semantic interpretation.

### The Blending Formula

The final **blended score** is computed using a weighted combination of the RRF score and the reranker score:

```

blendedScore = (RRF weight × RRF score) + ((1 - RRF weight) × reranker score)

```

Documents are then re-sorted by this blended score in descending order. The formula ensures that as the RRF rank decreases (lower confidence in the initial retrieval), the reranker gains more influence over the final ordering.

### Implementation in src/store.ts

The core logic for position-aware blending resides in [`src/store.ts`](https://github.com/tobi/qmd/blob/main/src/store.ts), where the weight selection and score calculation are implemented. The code determines the appropriate tier based on the document's RRF rank, retrieves the corresponding weight, and computes the blended score before final sorting.

## Practical Implementation

The following TypeScript example demonstrates the position-aware blending logic as implemented in QMD:

```typescript
// Assume `rrfResults` contains hybrid search results with:
//   - rrfScore: number   // RRF fusion score
//   - rrfRank: number    // 1-based position in RRF list
// `reranked` contains cross-encoder outputs with:
//   - score: number      // reranker confidence (0-1)
//   - docId: string

function blendScores(
  rrfResults: {docId: string; rrfScore: number; rrfRank: number}[],
  reranked: {docId: string; score: number}[]
) {
  const rerankMap = new Map(reranked.map(r => [r.docId, r.score]));

  return rrfResults.map(r => {
    // Position-aware weight selection
    const rrfWeight =
      r.rrfRank <= 3 ? 0.75 :
      r.rrfRank <= 10 ? 0.60 : 0.40;

    const rerankScore = rerankMap.get(r.docId) ?? 0;
    
    // Blended score calculation
    const blendedScore = 
      rrfWeight * r.rrfScore + 
      (1 - rrfWeight) * rerankScore;

    return {docId: r.docId, score: blendedScore};
  })
  .sort((a, b) => b.score - a.score);
}

```

When users execute `qmd query "<your question>"`, this blending logic automatically applies the tiered weighting to combine retrieval and neural signals without requiring manual configuration.

## Why Position-Aware Blending Matters

Traditional reranking approaches often allow the cross-encoder to completely override initial retrieval scores. This creates a vulnerability: when query expansion generates terms that don't appear in the original query, the neural model might demote documents that actually contain exact keyword matches.

**Position-aware blending** solves this by:

- **Preserving high-confidence lexical matches**: Top-ranked RRF results (ranks 1-3) retain 75% influence from the retrieval score, making them resistant to reranker override
- **Allowing semantic uplift**: Lower-ranked documents (rank 11+) give 60% weight to the reranker, enabling the neural model to surface semantically relevant content that sparse retrieval missed
- **Maintaining query fidelity**: The tiered approach specifically protects exact-match results when expanded queries introduce semantic drift

## Summary

- QMD implements **position-aware blending** to merge Reciprocal Rank Fusion (RRF) scores with cross-encoder reranker scores.
- The technique applies tiered weights based on RRF rank: **0.75 for ranks 1-3**, **0.60 for ranks 4-10**, and **0.40 for rank 11+**.
- The final score formula combines these weights: `blendedScore = (weight × RRF score) + ((1-weight) × reranker score)`.
- Core implementation resides in [`src/store.ts`](https://github.com/tobi/qmd/blob/main/src/store.ts), which handles weight selection and score computation.
- This approach preserves high-confidence lexical matches while allowing neural reranking to improve lower-confidence results.

## Frequently Asked Questions

### What is position-aware blending in QMD?

Position-aware blending is a scoring technique that combines initial retrieval scores (from RRF) with neural reranker scores using weights that vary based on the document's original rank. Higher-ranked documents retain more influence from the retrieval score, protecting strong lexical matches from being overridden by the cross-encoder.

### How does QMD calculate the final blended score?

QMD calculates the blended score using the formula: `blendedScore = (RRF weight × RRF score) + ((1 - RRF weight) × reranker score)`. The RRF weight is determined by the document's rank tier: 0.75 for ranks 1-3, 0.60 for ranks 4-10, and 0.40 for ranks 11 and above.

### Why does QMD use different weights for different rank tiers?

The tiered weighting system preserves high-confidence retrieval results while allowing semantic reranking to improve lower-ranked candidates. Top-ranked documents (1-3) receive high RRF weights (0.75) to prevent the reranker from demoting exact keyword matches, especially when query expansion introduces terms not present in the original query. Lower-ranked documents receive less protection, enabling the neural model to surface semantically relevant content that sparse retrieval missed.

### Where is the blending logic implemented in the QMD codebase?

The position-aware blending logic is implemented in [`src/store.ts`](https://github.com/tobi/qmd/blob/main/src/store.ts). This file contains the code that selects the appropriate RRF weight based on document rank and computes the final blended score by combining the RRF score with the cross-encoder reranker score.