How Claude Context Implements Hybrid Search: Combining BM25 and Dense Vector Retrieval

Claude Context implements hybrid search by running BM25-based sparse retrieval and dense vector similarity search in parallel on a Milvus collection, then merging results using Reciprocal Rank Fusion (RRF).

The zilliztech/claude-context repository provides a complete hybrid retrieval pipeline that unifies lexical and semantic search capabilities. By storing both dense embeddings and BM25-generated sparse vectors in a single Milvus collection, the system captures exact term matches alongside conceptual similarity. This article breaks down the implementation details, from collection schema design to the RRF reranking algorithm.

Architecture Overview: The Hybrid Collection

The foundation of Claude Context's hybrid search lies in a dual-field Milvus collection schema defined in packages/core/src/vectordb/milvus-vectordb.ts. The createHybridCollection method establishes two vector fields:

  • Dense vector field: Named vector with type FloatVector storing embedding representations
  • Sparse vector field: Named sparse_vector with type SparseFloatVector storing BM25 term frequency data

A critical component is the built-in BM25 function content_bm25_emb, which attaches to the content text field. This function automatically computes sparse representations during document insertion, requiring no manual tokenization or TF-IDF calculation from the caller.

// Simplified schema structure from createHybridCollection
const fields = [
  { name: "content", type: DataType.VarChar, max_length: 65535 },
  { name: "vector", type: DataType.FloatVector, dim: dimension },
  { name: "sparse_vector", type: DataType.SparseFloatVector }
];
// BM25 function automatically processes the 'content' field

Step-by-Step Implementation Flow

Step 1: Create the Hybrid Collection

During initialization, createHybridCollection (lines 667-702 in milvus-vectordb.ts) configures the collection with both vector types and the BM25 function. This setup enables Milvus to handle sparse vector generation transparently when raw text is inserted.

Step 2: Index Documents with Dual Representations

The insertHybrid method (lines 593-616) handles document ingestion. Callers provide the dense embedding via the vector field and the original text via content. Milvus automatically invokes content_bm25_emb to populate sparse_vector, creating a complete hybrid index without duplicate storage logic.

await db.insertHybrid(collectionName, {
  content: "function validateUser() { ... }",
  vector: [0.023, -0.156, ...], // Dense embedding from your model
  // sparse_vector is auto-generated by BM25 function
});

Step 3: Build Parallel Search Requests

When semanticSearch is invoked with HYBRID_MODE enabled (the default), the system constructs two HybridSearchRequest objects in packages/core/src/context.ts (lines 440-452):

Dense Request:

  • Field: vector
  • Data: Query embedding vector
  • Parameters: nprobe: 10 for IVF index probing

Sparse Request:

  • Field: sparse_vector
  • Data: Raw query string (e.g., "find all functions that validate user input")
  • Parameters: drop_ratio_search: 0.2 to filter insignificant terms

Step 4: Execute Hybrid Search with RRF Fusion

The hybridSearch method (lines 618-675 in milvus-vectordb.ts) passes both requests to MilvusClient.search as a combined payload. Milvus executes the dense ANN search and BM25-driven sparse search simultaneously. The SDK then applies Reciprocal Rank Fusion with k: 100 to merge results:

const hybridResults = await db.hybridSearch(
  collectionName,
  [denseReq, sparseReq],
  {
    rerank: {
      strategy: 'rrf',
      params: { k: 100 }
    }
  }
);

RRF calculates a fused score by summing the reciprocal ranks from both result sets, eliminating the need for learned scoring models or complex calibration between sparse and dense similarity metrics.

Step 5: Map to Unified Results

Finally, results are mapped to HybridSearchResult objects and converted to SemanticSearchResult format in context.ts (lines 760-784). Each result contains the document chunk, file location, and the merged RRF relevance score.

Practical Implementation Examples

Enabling Hybrid Mode

Hybrid search is controlled by the HYBRID_MODE environment variable, checked via getIsHybrid in context.ts (lines 221-235):

import { envManager } from '@zilliz/claude-context-core';

const isHybrid = envManager.get('HYBRID_MODE')?.toLowerCase() !== 'false';
// Defaults to true; set HYBRID_MODE=false to disable

Indexing a Codebase

The high-level indexCodebase method orchestrates hybrid collection creation:

const codebasePath = '/path/to/project';
await context.indexCodebase(codebasePath);
// Internally calls createHybridCollection() and insertHybrid()

Applications interact with the hybrid pipeline through the standard semanticSearch API:

const query = 'authenticate user session';
const results = await context.semanticSearch(
  codebasePath,
  query,
  topK = 5,
  threshold = 0.3
);

results.forEach(r => {
  console.log(`[${r.score.toFixed(2)}] ${r.relativePath}:${r.startLine}`);
});

When hybrid mode is active, this automatically routes through the dual-request RRF pipeline.

Low-Level Hybrid API Access

For advanced use cases, access the Milus-specific implementation directly:

import { MilvusVectorDatabase } from '@zilliz/claude-context-core';

const db = new MilvusVectorDatabase({ address: 'localhost:19530' });

const denseReq = {
  data: queryEmbedding,
  anns_field: 'vector',
  param: { nprobe: 10 },
  limit: 10
};

const sparseReq = {
  data: 'raw query text',
  anns_field: 'sparse_vector',
  param: { drop_ratio_search: 0.2 },
  limit: 10
};

const results = await db.hybridSearch(
  'my_collection',
  [denseReq, sparseReq],
  { rerank: { strategy: 'rrf', params: { k: 100 } } }
);

Key Technical Components

Summary

  • Dual-schema design: Milvus collections store both FloatVector (dense) and SparseFloatVector (BM25) fields, linked by the content_bm25_emb function
  • Parallel execution: Dense and sparse searches run simultaneously via separate HybridSearchRequest configurations
  • RRF fusion: Results merge using Reciprocal Rank Fusion (k: 100) without requiring trained cross-encoders
  • Transparent API: The semanticSearch method automatically handles hybrid logic when HYBRID_MODE is enabled (default)
  • Lexical + Semantic: BM25 captures exact identifier matches while dense vectors handle semantic similarity and paraphrasing

Frequently Asked Questions

What is Reciprocal Rank Fusion (RRF) in Claude Context?

RRF is a rank aggregation algorithm that combines result lists from multiple retrieval methods without requiring relevance scores to be calibrated or normalized. Claude Context uses RRF with k: 100, calculating a fused score as the sum of reciprocals: 1/(k + rank). This gives higher weights to documents that appear near the top of both the BM25 and dense vector result lists, effectively boosting items that excel in both lexical and semantic relevance.

How does the BM25 function generate sparse vectors automatically?

The content_bm25_emb function attaches to the content text field during collection creation in createHybridCollection. When documents are inserted via insertHybrid, Milvus executes this built-in function to compute term frequencies and inverse document frequencies on-the-fly. The caller only needs to provide the raw text and dense embedding; Milvus handles tokenization and sparse vector generation internally, storing the result in the sparse_vector field.

Can I disable hybrid search and use only dense vector retrieval?

Yes. Set the environment variable HYBRID_MODE=false in your .env file or environment configuration. The getIsHybrid helper in context.ts checks this variable, and when disabled, the semanticSearch method bypasses the parallel request construction and RRF fusion, falling back to standard dense vector search against the vector field only.

What configuration parameters tune the hybrid search behavior?

Two key parameters control the search behavior: nprobe: 10 for the dense request determines how many clusters to probe in the IVF index, trading recall for speed, while drop_ratio_search: 0.2 for the sparse request filters out less significant terms from the BM25 query vector, reducing noise in the lexical matching. The RRF constant k: 100 can also be adjusted in the hybridSearch options to control how aggressively rank positions influence the final fused score.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →