How RRF (Reciprocal Rank Fusion) Reranking Works in Claude Context's Hybrid Search
Claude Context combines dense vector and sparse lexical search results using Reciprocal Rank Fusion (RRF), which calculates composite relevance scores based on document rankings across both search modalities rather than raw similarity scores.
Claude Context, an open-source retrieval system developed by Zilliz, implements a two-stage hybrid search that leverages Milvus's native RRF capabilities. The Context.search method orchestrates parallel semantic and lexical queries, then applies reciprocal rank fusion to produce a unified, reranked result set. This approach eliminates the complexity of score normalization while balancing deep semantic understanding with precise keyword matching.
Constructing the Hybrid Search Request with RRF Configuration
The hybrid search process begins in packages/core/src/context.ts, where the Context.search method constructs two distinct HybridSearchRequest objects. These represent the dense vector embedding query and the sparse lexical query respectively.
According to lines 58-66 of packages/core/src/context.ts, the implementation builds the search requests as follows:
const searchRequests: HybridSearchRequest[] = [
{ data: queryEmbedding.vector, anns_field: "vector", param: { "nprobe": 10 }, limit: topK },
{ data: query, anns_field: "sparse_vector", param: { "drop_ratio_search": 0.2 }, limit: topK }
];
The RRF reranking strategy is configured via the rerank option passed to the hybridSearch call:
const searchResults = await this.vectorDatabase.hybridSearch(
collectionName,
searchRequests,
{
rerank: { strategy: 'rrf', params: { k: 100 } },
limit: topK,
filterExpr
}
);
This configuration directs Milvus to fuse the two result sets using the Reciprocal Rank Fusion algorithm rather than returning raw concatenated results.
Propagating RRF Parameters to Milvus
The reranking configuration flows from the high-level Context API down to the underlying Milvus client in packages/core/src/vectordb/milvus-vectordb.ts. As shown in lines 45-51, the wrapper packages the RRF parameters into Milvus's expected format:
const rerank_strategy = {
strategy: "rrf",
params: { k: 100 }
};
The MilvusVectorDB.hybridSearch method forwards this object directly to Milvus's hybrid search API, ensuring the database engine handles the computational work of result fusion. This architecture keeps the application layer lightweight while delegating the ranking computation to the optimized database engine.
The RRF Algorithm and Scoring Mechanism
Milvus implements the standard Reciprocal Rank Fusion algorithm to combine the two ranked lists. For each document appearing in either the dense or sparse results, the system calculates a composite score using the formula:
score_RRF = sum(1 / (k + rank_i))
Where:
- rank_i represents the position of the document in the i-th sub-search result list
- k is a constant smoothing parameter (defaulting to 100 in Claude Context)
- The summation runs across both sub-searches (dense and sparse)
The k parameter serves a critical smoothing function. By defaulting to 100, Claude Context ensures that top-ranked documents receive high scores while lower-ranked items still contribute meaningfully to the fusion. This prevents the system from over-penalizing documents that rank moderately well in one modality but strongly in another.
Implementing RRF Reranking in Practice
Using the Context.search API
For most applications, the high-level Context.search method abstracts the complexity of RRF configuration. The following TypeScript example demonstrates a hybrid search with custom RRF parameters:
import { Context } from '@zilliz/claude-context';
const query = "How does RRF work in Claude Context?";
const results = await context.search(query, {
hybrid: true,
topK: 10,
rerank: { strategy: 'rrf', params: { k: 150 } }
});
results.forEach(r => {
console.log(`[${r.score.toFixed(4)}] ${r.relativePath}:${r.startLine}-${r.endLine}`);
console.log(r.content);
});
The implementation automatically generates the dual search requests and attaches the RRF configuration before dispatching to the vector database layer.
Direct Milvus Wrapper Access
For scenarios requiring granular control over search parameters, you can interact directly with the MilvusVectorDB class:
import { MilvusVectorDB } from '@zilliz/claude-context/vectordb';
const db = new MilvusVectorDB(/* configuration */);
const denseReq = {
data: denseVector,
anns_field: 'vector',
param: { nprobe: 10 },
limit: 10
};
const sparseReq = {
data: textQuery,
anns_field: 'sparse_vector',
param: { drop_ratio_search: 0.2 },
limit: 10
};
const hybridResults = await db.hybridSearch(
'my_collection',
[denseReq, sparseReq],
{
rerank: { strategy: 'rrf', params: { k: 100 } },
limit: 10
}
);
This approach exposes the underlying anns_field parameters, allowing precise tuning of the dense vector probe count (nprobe) and sparse vector drop ratio (drop_ratio_search).
Performance Advantages of RRF in Hybrid Search
Claude Context selects RRF over alternative fusion methods for three specific technical reasons:
- Computational efficiency: RRF requires no additional model inference or embedding calculations, operating solely on result rankings.
- Scale independence: Because RRF uses rank positions rather than raw similarity scores, it inherently handles the scale mismatches between dense vector cosine similarities and sparse vector term frequencies.
- Configurable smoothing: The
kparameter provides explicit control over the influence of lower-ranked documents, with the default value of 100 offering balanced weighting between precision and recall.
Summary
- Claude Context implements Reciprocal Rank Fusion in
packages/core/src/context.tsby constructing dualHybridSearchRequestobjects for dense and sparse vectors. - The RRF configuration propagates through
packages/core/src/vectordb/milvus-vectordb.ts(lines 45-51) to Milvus with a default smoothing parameter of k=100. - The algorithm calculates document scores as the sum of reciprocal ranks across both search modalities using the formula
1/(k + rank). - RRF eliminates the need for score normalization while remaining computationally lightweight compared to learned reranking models.
- Developers can customize the
kparameter via thererankoption in both high-levelContext.searchcalls and low-levelMilvusVectorDB.hybridSearchinvocations.
Frequently Asked Questions
What is the default k parameter value in Claude Context's RRF implementation?
The default value is 100, as specified in the rerank configuration object passed to Milvus. This value appears in packages/core/src/context.ts and packages/core/src/vectordb/milvus-vectordb.ts. The k parameter smooths the contribution of lower-ranked documents, ensuring that a document ranking highly in either the dense or sparse search receives a strong combined score.
How does RRF handle score scale differences between dense and sparse searches?
RRF operates exclusively on rank positions rather than raw similarity scores. Because the algorithm uses the formula 1/(k + rank), it is invariant to the absolute values produced by dense vector embeddings (typically cosine similarities) and sparse lexical scores (typically BM25 or TF-IDF variants). This robustness eliminates the need for complex score calibration or normalization between the two search modalities.
Can I customize the RRF k parameter when using the Context.search method?
Yes, the k parameter is configurable through the rerank options object. When calling context.search(), pass a custom value via rerank: { strategy: 'rrf', params: { k: 150 } } to adjust the smoothing behavior. Lower values emphasize top-ranked results more aggressively, while higher values allow deeper-ranked items to contribute more significantly to the final fusion score.
Which Milvus collection fields are used for the hybrid search sub-queries?
Claude Context queries two specific fields defined in the search requests within packages/core/src/context.ts: the dense vector field (anns_field: "vector") and the sparse vector field (anns_field: "sparse_vector"). These correspond to the semantic embedding generated by the LLM and the lexical sparse representation produced by the text splitter, respectively.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →