WeKnora Hybrid Retrieval Pipeline: Combining BM25, Dense Vectors, GraphRAG, and Parent-Child Chunking
WeKnora's hybrid retrieval pipeline concurrently executes BM25 sparse retrieval, BM2S dense vector search with GraphRAG knowledge graph traversal, and parent-child chunk enrichment to deliver high-precision, contextually grounded answers.
Tencent's WeKnora implements an advanced hybrid retrieval pipeline that addresses the limitations of single-method retrieval by fusing complementary search strategies. This architecture integrates traditional inverted-index techniques with modern semantic embeddings and hierarchical document structures, enabling robust retrieval across diverse query types. The pipeline is model-agnostic, supporting interchangeable embedding models and rerankers while maintaining a consistent interface through the MCP server implementation.
Core Components of the Hybrid Architecture
The retrieval system combines three distinct technologies that operate in concert during query execution. Each component addresses specific failure modes of the others, creating a resilient search layer.
BM25 Sparse Retrieval
BM25 provides the foundation of the hybrid pipeline through classic keyword matching against inverted-index data. As implemented in the retrieval orchestration layer, this component delivers a fast first-pass filter that captures exact-term hits and rare technical terminology. The sparse retrieval stage generates an initial candidate shortlist thatanchors the subsequent semantic search phases to relevant document subsets.
BM2S Dense Vector Search with GraphRAG
The BM2S (dense retrieval) component utilizes embedding models such as BGE or OpenAI embeddings to identify semantically similar chunks beyond lexical matches. This dense search integrates with GraphRAG, which traverses knowledge graph relationships built from chunk metadata. Specifically, the graph layer expands queries by following parent-child links and semantic edges between chunks, then applies a reranker to the expanded candidate set. According to the source code in weknora_mcp_server.py, this combination runs concurrently with BM25 to maximize recall while maintaining precision.
Parent-Child Chunking Hierarchy
Documents undergo hierarchical splitting during ingestion, creating child chunks (fine-grained passages) linked to parent chunks (broader sections or pages). The splitter utilities in docreader/splitter/splitter.py establish these relationships during parsing, while docreader/parser/base_parser.py attaches parent IDs to chunk metadata. This hierarchy is stored alongside vector embeddings, enabling the retrieval layer to fetch entire logical blocks when any child chunk matches. When a child chunk is retrieved, the system automatically resolves its parent context, ensuring the LLM receives both granular evidence and surrounding narrative flow.
Pipeline Execution Flow
The hybrid orchestration logic resides primarily in mcp-server/weknora_mcp_server.py, specifically within the chat tool implementation. The execution follows a concurrent retrieval and enrichment pattern:
-
Query Reception: The client assembles a RAG pipeline request via
client.chat(), specifying knowledge base IDs and query parameters. -
Concurrent Hybrid Search: The backend simultaneously executes BM25 sparse retrieval and BM2S dense-vector search with GraphRAG traversal. The graph component expands semantic neighborhoods by traversing parent-child relationships identified during document parsing.
-
Chunk Enrichment: Retrieved child chunks are automatically expanded with their parent chunks (and optionally sibling chunks) using the hierarchy stored in the vector metadata. This step resolves the parent IDs attached during the parsing stage in
docreader/parser/base_parser.py. -
Merging and Deduplication: Results from BM25 and BM2S/GraphRAG sources are merged, re-ranked, and de-duplicated based on semantic similarity and graph distance metrics.
-
LLM Generation: The final enriched chunk set is fed to the language model for answer synthesis, combining precise keyword matches with broad semantic context.
Implementation Details and Code Examples
The pipeline exposes both high-level chat interfaces and low-level retrieval functions. The hybrid_retrieve helper wraps the complexity of coordinating BM25, dense search, and parent-child expansion into a single call.
Invoking the Complete Pipeline
from weknora_mcp_server import client
# Create a chat session bound to a knowledge base
session = client.create_session(kb_id="my-docs", max_rounds=3)
# Execute hybrid retrieval and generation
answer = client.chat(
session_id=session["id"],
query="How does WeKnora generate a knowledge graph from uploaded PDFs?",
knowledge_base_ids=["my-docs"],
web_search_enabled=False,
)
print(answer["content"])
Direct Hybrid Retrieval Access
from weknora_mcp_server import client
# Manually trigger hybrid search for debugging or custom workflows
results = client.hybrid_retrieve(
query="What is parent-child chunking?",
kb_id="my-docs",
top_k=10,
)
for r in results:
print(f"Score: {r['score']:.2f} Chunk: {r['text'][:120]}…")
Key Source Files
The hybrid retrieval implementation spans several critical modules:
-
mcp-server/weknora_mcp_server.py: Contains thechattool entry point that orchestrates the full RAG pipeline, coordinating between sparse and dense retrieval paths. -
docreader/splitter/splitter.py: Implements the hierarchical splitting logic that creates parent-child chunk relationships during document ingestion. -
docreader/parser/base_parser.py: Attaches metadata including parent IDs to chunks, establishing the relationships later traversed by GraphRAG. -
docreader/utils/request.py: Handles vector store queries for the BM2S dense retrieval component and manages the combination of dense results with BM25 outputs.
Summary
-
Triple-Method Fusion: WeKnora combines BM25 (keyword precision), BM2S dense vectors (semantic breadth), and GraphRAG (relational context) to maximize retrieval accuracy across query types.
-
Hierarchical Context: Parent-child chunking ensures retrieved passages include both specific evidence and broader document context, reducing information fragmentation.
-
Concurrent Execution: The pipeline runs BM25 and dense/GraphRAG searches simultaneously, merging results through a deduplication and reranking layer before LLM generation.
-
Model Agnostic Design: Embedding models and rerankers can be swapped via configuration without modifying the core pipeline logic in
weknora_mcp_server.py. -
Graph-Augmented Retrieval: The knowledge graph is built from parent-child links and chunk relationships, enabling traversal-based result expansion beyond vector similarity alone.
Frequently Asked Questions
What is the difference between BM25 and BM2S in WeKnora's pipeline?
BM25 performs sparse, keyword-based retrieval using inverted indices to find exact term matches, while BM2S refers to the dense vector search component that uses neural embeddings to identify semantically similar content. BM2S operates alongside GraphRAG to traverse knowledge graph relationships, whereas BM25 provides the initial filtering layer based on lexical overlap.
How does parent-child chunking improve retrieval quality?
Parent-child chunking creates a hierarchical document structure where fine-grained child chunks link to broader parent sections. When the system retrieves a child chunk via BM25 or dense search, it automatically includes the parent chunk's context, ensuring the LLM receives both specific evidence and surrounding narrative. This prevents the "lost context" problem common in flat chunking strategies.
What role does GraphRAG play in the hybrid retrieval process?
GraphRAG enhances the BM2S dense retrieval by treating chunks as nodes in a knowledge graph connected by parent-child and semantic relationships. After initial vector similarity identifies candidate chunks, GraphRAG traverses these relationships to expand the result set with logically connected passages, then reranks the combined pool to surface the most relevant context chains.
How can developers customize the embedding models in this pipeline?
The hybrid retrieval pipeline is model-agnostic, allowing configuration of any embedding model (such as BGE or OpenAI embeddings) through the model configuration files without modifying the core retrieval logic in weknora_mcp_server.py. The docreader/utils/request.py module handles the abstraction between the configured model and the vector store queries.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →