How the Multi-Phase Retrieval Pipeline Works in LLM Wiki: A 4-Stage Architecture
LLM Wiki answers queries using a four-stage retrieval pipeline that combines tokenized keyword search, optional vector semantic search, graph-based expansion, and token-budget-aware context assembly to maximize recall while respecting LLM context limits.
The multi-phase retrieval pipeline in LLM Wiki progressively expands the candidate document set to ensure high recall before intelligently selecting content for the LLM. Implemented in the open-source nashsu/llm_wiki repository, this architecture balances computational efficiency with semantic richness by layering fast token matching, vector similarity search, and knowledge graph traversal. The system achieves approximately 71% recall compared to 58% for keyword-only approaches, while strictly enforcing configurable context window limits ranging from 4K to 1M tokens.
Phase 1: Tokenized Search (Fast Keyword Matching)
The pipeline begins with a high-performance language-aware keyword search implemented in src/lib/search.ts. The tokenizeQuery() function (lines 38-59) processes queries differently based on language: English text is split on whitespace with stop-words removed, while Chinese queries are segmented into overlapping bigrams plus individual characters to ensure partial matching.
This phase searches both the wiki/ directory and the raw source folder (raw/sources/), applying a fixed +10 relevance bonus to page titles. The function calls the Rust backend command search_project via Tauri, which returns an initial ranked list of candidates based on token overlap.
Phase 1.5: Optional Vector Semantic Search
When vector search is enabled via embeddingConfig from useWikiStore.getState().embeddingConfig in src/lib/search.ts, the pipeline generates query embeddings using any OpenAI-compatible /v1/embeddings endpoint. These embeddings are compared against a LanceDB vector index maintained in the Rust backend using cosine similarity.
Vector hits are merged with the token search results, boosting the scores of documents that match both criteria while introducing semantically related pages that lack keyword overlap. This optional phase significantly improves recall by capturing conceptual relationships beyond lexical matching.
Phase 2: Graph Expansion (Relationship Discovery)
The top-K results from the retrieval phases become seed nodes for a 2-hop graph traversal implemented in src/lib/wiki-graph.ts and src/lib/graph-relevance.ts. The expandGraph() routine applies a 4-signal relevance model that scores neighbor pages based on direct wikilinks, source file overlap, Adamic-Adar index, and type affinity.
This phase discovers highly relevant documents that might not contain the original query keywords but are strongly connected to the initial results through the knowledge graph structure. By traversing relationships with decay scoring, Phase 2 effectively expands the candidate set based on topological relevance.
Phase 3: Budget-Controlled Context Assembly
Before invoking the LLM, the pipeline enforces strict token budget constraints configured in src/stores/wiki-store.ts. The system allocates the configurable context window (4K to 1M tokens) proportionally: approximately 60% to wiki page content, 20% to recent chat history, 5% to index.md, and 15% to system prompts.
Pages are ranked by a combined score integrating keyword/vector relevance and graph relevance scores. The highest-scoring pages are included in descending order until the allocated token budget for wiki content is exhausted, ensuring optimal use of the limited context window.
Phase 4: Final Context Assembly and LLM Invocation
In the final stage, selected pages are concatenated with their full content rather than summaries. The system prompt constructed in src/lib/prompt.ts injects purpose.md, language-specific rules, citation formatting instructions, and index.md metadata.
Each page is assigned a numbered reference (e.g., [1], [2]) that the LLM uses to cite sources in its response. The complete payload is sent to the Rust backend via invoke("chat"), which manages the actual LLM API call and returns the generated answer with inline citations.
Implementation Examples
The following TypeScript examples demonstrate how to interact with the retrieval pipeline from client components:
// Example: performing a search from a component
import { searchWiki } from "@/lib/search";
async function runQuery(projectPath: string, query: string) {
// Phase 1 + optional Phase 1.5 (handled internally)
const results = await searchWiki(projectPath, query);
// The returned results already contain the combined
// token + vector scores and are ready for graph expansion.
console.log("Top results:", results.slice(0, 5));
}
// Example: tokenizing a query (used in Phase 1)
import { tokenizeQuery } from "@/lib/search";
const tokens = tokenizeQuery("How does LLM Wiki retrieve information?");
console.log(tokens);
// → ['how', 'does', 'llm', 'wiki', 'retrieve', 'information']
Key Source Files
Understanding the pipeline requires familiarity with these specific modules:
src/lib/search.ts– ImplementstokenizeQuery()andsearchWiki(), the primary entry points for Phase 1 and 1.5.src/lib/wiki-graph.ts– Provides graph data structures and the 2-hop expansion algorithm used in Phase 2.src/lib/graph-relevance.ts– Defines the 4-signal relevance scoring model (wikilinks, source overlap, Adamic-Adar, type affinity).src/stores/wiki-store.ts– ManagesembeddingConfig,contextWindowsettings, and token budget calculations for Phase 3.src/lib/prompt.ts– Constructs the system prompt template with citations and metadata for Phase 4.src-tauri/src/api_server.rs– Rust backend hostingsearch_projectand LanceDB vector operations.
Summary
- Four-stage architecture combines tokenized search, optional vector retrieval, graph expansion, and budget-aware assembly to maximize recall while controlling token usage.
- Language-aware tokenization handles English stop-word removal and Chinese bigram segmentation in
src/lib/search.ts. - Graph relevance model uses four signals (wikilinks, source overlap, Adamic-Adar, type affinity) during 2-hop traversal to discover related content.
- Strict token budgeting allocates context window space proportionally across wiki content, chat history, and system prompts.
- Traceable citations via numbered references ensure answers can be verified against source documents.
Frequently Asked Questions
How does LLM Wiki handle multilingual queries during tokenization?
The tokenizeQuery() function in src/lib/search.ts detects language characteristics and applies distinct strategies: English queries are split on whitespace with common stop-words filtered out, while Chinese text is segmented into overlapping bigrams plus individual characters. This ensures effective partial matching regardless of whether the query uses space-delimited or continuous script.
What is the purpose of the graph expansion phase?
Graph expansion discovers relevant documents that lack keyword or vector similarity to the original query but are topologically close to high-scoring seed documents in the knowledge graph. By traversing 2 hops from seed nodes using the 4-signal relevance model, Phase 2 improves recall for conceptually related content that might otherwise be missed by surface-level matching.
How does the system prevent exceeding the LLM's context window?
Phase 3 implements a configurable token budget system in src/stores/wiki-store.ts that allocates percentages of the total context window (ranging from 4K to 1M tokens) to specific content types. Pages are ranked by combined relevance scores and added to the context only until the wiki content allocation (approximately 60%) is exhausted, ensuring rigid adherence to the limit.
Can vector search be disabled to improve search latency?
Yes, vector semantic search is optional and controlled by the embeddingConfig setting read from useWikiStore.getState().embeddingConfig. When disabled, the pipeline skips Phase 1.5 and proceeds directly from tokenized search to graph expansion, reducing latency at the cost of semantic recall (approximately 58% recall versus 71% with vectors enabled).
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →