The Four Phases of the Multi-Phase Retrieval Pipeline in LLM Wiki
LLM Wiki implements a four-phase retrieval pipeline comprising Tokenized Search, Graph Expansion, Budget Control, and Context Assembly to deliver high-quality, token-budget-aware context for large language models.
The nashsu/llm_wiki repository enhances traditional retrieval-augmented generation (RAG) through a sophisticated multi-phase retrieval pipeline that blends lexical, semantic, and graph-based signals. This architecture ensures that the assembled context remains within configurable token limits while maximizing information relevance.
Phase 1 – Tokenized Search
The pipeline begins with fast keyword lookup across both wiki and raw source collections. Implemented in src/lib/search.ts (lines 38-60), this phase handles multilingual tokenization through distinct strategies.
English tokenization splits content on whitespace and punctuation while filtering out stop-words to generate searchable terms. Chinese tokenization employs CJK bigram splitting combined with character-level tokens to capture semantic units in logographic scripts. The system applies a title-match bonus of +10 to boost exact-title hits, ensuring high-precision matches surface immediately.
Optional Sub-Phase 1.5 – Vector Semantic Search
When enabled in Settings, the pipeline inserts an embedding-based ANN search between Phase 1 and Phase 2. Located in src/lib/embedding.ts, this optional step generates vectors via an OpenAI-compatible /v1/embeddings endpoint. Results merge with the tokenized search output and undergo re-ranking before proceeding to graph expansion.
Phase 2 – Graph Expansion
Following initial retrieval, the system broadens the result set by exploiting the internal knowledge graph. The top results from Phase 1 serve as seed nodes for traversal.
The expansion logic, housed in src/lib/graph-search.ts, employs a 4-signal relevance model that evaluates link weight, type affinity, and other relational metrics. A 2-hop traversal with decay discovers related concepts beyond the initial keyword hits, ensuring the LLM receives contextualized information that purely lexical search might miss.
Phase 3 – Budget Control
This phase guarantees that assembled context respects the LLM's token window. Implemented in src/lib/context-budget.ts, the system supports configurable context windows ranging from 4K to 1M tokens.
Tokens allocate proportionally across four categories: approximately 60% for wiki pages, 20% for chat history, 5% for the index, and 15% for system instructions. Pages rank by a combined search-plus-graph relevance score and trim to fit the budget, preventing context overflow while preserving the most pertinent information.
Phase 4 – Context Assembly
The final stage packages selected pages into the prompt sent to the LLM. Unlike systems that provide excerpts, LLM Wiki delivers numbered pages with full content to maximize reasoning depth.
The system prompt concatenates essential project files including purpose.md, language-specific rules, citation format definitions, and index.md. The LLM receives explicit instructions to cite pages using bracketed numbers ([1], [2], etc.), creating a traceable audit trail for every generated claim.
Implementation Example
The following TypeScript snippet demonstrates how a frontend component invokes the complete pipeline:
import { searchWiki } from '@/lib/search'
import { useWikiStore } from '@/stores/wiki-store'
// Phase 1: Tokenized (and optional vector) search
const query = 'distributed consensus algorithms'
const projectPath = '/path/to/your/wiki'
const results = await searchWiki(projectPath, query)
// Phase 2: Graph expansion occurs internally when requesting full context
// Phase 3: Configure token budget via the store
useWikiStore.getState().setContextBudget({
maxTokens: 200_000,
allocation: { wiki: 0.6, chat: 0.2, index: 0.05, system: 0.15 },
})
// Phase 4: Assemble the final prompt
const { assembledPrompt, citedPages } = await invoke('assemble_context', {
projectPath,
query,
results,
})
// Send to LLM
const completion = await fetch('/api/chat', {
method: 'POST',
body: JSON.stringify({
messages: [{ role: 'user', content: assembledPrompt }]
}),
})
This flow demonstrates how raw queries progress through tokenization, optional semantic enrichment, graph-based expansion, budget-constrained trimming, and final assembly into a citation-ready prompt.
Summary
- Tokenized Search (
search.ts): Performs multilingual keyword lookup with title-match boosting. - Graph Expansion (
graph-search.ts): Extends results via 2-hop graph traversal using a 4-signal relevance model. - Budget Control (
context-budget.ts): Enforces configurable token limits with proportional allocation across wiki, chat, index, and system content. - Context Assembly: Delivers numbered full-page content and structured system prompts with citation support.
Frequently Asked Questions
What is the optional Phase 1.5 in the LLM Wiki retrieval pipeline?
Phase 1.5 is an optional vector semantic search that runs between tokenized search and graph expansion. When enabled, it generates embeddings through an OpenAI-compatible API, performs approximate nearest neighbor (ANN) search, and merges results with Phase 1 output before re-ranking. This sub-phase does not constitute a main phase but enhances recall for semantic queries that keyword search might miss.
How does the graph expansion phase improve retrieval quality?
Graph Expansion (Phase 2) improves quality by traversing relationships beyond keyword matches. It uses the top Phase 1 results as seed nodes and applies a 4-signal relevance model to score neighboring pages during a 2-hop traversal. This captures related concepts and implicit connections within the knowledge graph, providing the LLM with broader contextual awareness than lexical search alone.
What token budget allocations does LLM Wiki support?
The Budget Control phase supports context windows from 4,000 to 1,000,000 tokens. The default allocation distributes approximately 60% to wiki content, 20% to chat history, 5% to the index, and 15% to system instructions. Users can configure these percentages through the setContextBudget API to prioritize specific content types based on their use case.
How are sources cited in the assembled context?
During Context Assembly (Phase 4), the system assigns sequential numbers to selected pages and instructs the LLM to cite sources using bracketed identifiers like [1] or [2]. The assembled prompt includes full page content (not excerpts) alongside system files such as purpose.md and index.md, ensuring the model can reference specific documentation while generating responses.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →