# The Four Phases of the Multi-Phase Retrieval Pipeline in LLM Wiki

> Discover the four phases of LLM Wiki's multi-phase retrieval pipeline: Tokenized Search, Graph Expansion, Budget Control, and Context Assembly for efficient context delivery.

- Repository: [nash_su/llm_wiki](https://github.com/nashsu/llm_wiki)
- Tags: deep-dive
- Published: 2026-09-12

---

**LLM Wiki implements a four-phase retrieval pipeline comprising Tokenized Search, Graph Expansion, Budget Control, and Context Assembly to deliver high-quality, token-budget-aware context for large language models.**

The `nashsu/llm_wiki` repository enhances traditional retrieval-augmented generation (RAG) through a sophisticated multi-phase retrieval pipeline that blends lexical, semantic, and graph-based signals. This architecture ensures that the assembled context remains within configurable token limits while maximizing information relevance.

## Phase 1 – Tokenized Search

The pipeline begins with fast keyword lookup across both wiki and raw source collections. Implemented in [`src/lib/search.ts`](https://github.com/nashsu/llm_wiki/blob/main/src/lib/search.ts) (lines 38-60), this phase handles multilingual tokenization through distinct strategies.

**English tokenization** splits content on whitespace and punctuation while filtering out stop-words to generate searchable terms. **Chinese tokenization** employs CJK bigram splitting combined with character-level tokens to capture semantic units in logographic scripts. The system applies a **title-match bonus** of +10 to boost exact-title hits, ensuring high-precision matches surface immediately.

### Optional Sub-Phase 1.5 – Vector Semantic Search

When enabled in Settings, the pipeline inserts an embedding-based ANN search between Phase 1 and Phase 2. Located in [`src/lib/embedding.ts`](https://github.com/nashsu/llm_wiki/blob/main/src/lib/embedding.ts), this optional step generates vectors via an OpenAI-compatible `/v1/embeddings` endpoint. Results merge with the tokenized search output and undergo re-ranking before proceeding to graph expansion.

## Phase 2 – Graph Expansion

Following initial retrieval, the system broadens the result set by exploiting the internal knowledge graph. The top results from Phase 1 serve as **seed nodes** for traversal.

The expansion logic, housed in [`src/lib/graph-search.ts`](https://github.com/nashsu/llm_wiki/blob/main/src/lib/graph-search.ts), employs a **4-signal relevance model** that evaluates link weight, type affinity, and other relational metrics. A **2-hop traversal with decay** discovers related concepts beyond the initial keyword hits, ensuring the LLM receives contextualized information that purely lexical search might miss.

## Phase 3 – Budget Control

This phase guarantees that assembled context respects the LLM's token window. Implemented in [`src/lib/context-budget.ts`](https://github.com/nashsu/llm_wiki/blob/main/src/lib/context-budget.ts), the system supports configurable context windows ranging from **4K to 1M tokens**.

Tokens allocate proportionally across four categories: approximately **60%** for wiki pages, **20%** for chat history, **5%** for the index, and **15%** for system instructions. Pages rank by a combined search-plus-graph relevance score and trim to fit the budget, preventing context overflow while preserving the most pertinent information.

## Phase 4 – Context Assembly

The final stage packages selected pages into the prompt sent to the LLM. Unlike systems that provide excerpts, LLM Wiki delivers **numbered pages with full content** to maximize reasoning depth.

The system prompt concatenates essential project files including [`purpose.md`](https://github.com/nashsu/llm_wiki/blob/main/purpose.md), language-specific rules, citation format definitions, and [`index.md`](https://github.com/nashsu/llm_wiki/blob/main/index.md). The LLM receives explicit instructions to cite pages using bracketed numbers (`[1]`, `[2]`, etc.), creating a traceable audit trail for every generated claim.

## Implementation Example

The following TypeScript snippet demonstrates how a frontend component invokes the complete pipeline:

```typescript
import { searchWiki } from '@/lib/search'
import { useWikiStore } from '@/stores/wiki-store'

// Phase 1: Tokenized (and optional vector) search
const query = 'distributed consensus algorithms'
const projectPath = '/path/to/your/wiki'
const results = await searchWiki(projectPath, query)

// Phase 2: Graph expansion occurs internally when requesting full context
// Phase 3: Configure token budget via the store
useWikiStore.getState().setContextBudget({
  maxTokens: 200_000,
  allocation: { wiki: 0.6, chat: 0.2, index: 0.05, system: 0.15 },
})

// Phase 4: Assemble the final prompt
const { assembledPrompt, citedPages } = await invoke('assemble_context', {
  projectPath,
  query,
  results,
})

// Send to LLM
const completion = await fetch('/api/chat', {
  method: 'POST',
  body: JSON.stringify({ 
    messages: [{ role: 'user', content: assembledPrompt }] 
  }),
})

```

This flow demonstrates how raw queries progress through tokenization, optional semantic enrichment, graph-based expansion, budget-constrained trimming, and final assembly into a citation-ready prompt.

## Summary

- **Tokenized Search** ([`search.ts`](https://github.com/nashsu/llm_wiki/blob/main/search.ts)): Performs multilingual keyword lookup with title-match boosting.
- **Graph Expansion** ([`graph-search.ts`](https://github.com/nashsu/llm_wiki/blob/main/graph-search.ts)): Extends results via 2-hop graph traversal using a 4-signal relevance model.
- **Budget Control** ([`context-budget.ts`](https://github.com/nashsu/llm_wiki/blob/main/context-budget.ts)): Enforces configurable token limits with proportional allocation across wiki, chat, index, and system content.
- **Context Assembly**: Delivers numbered full-page content and structured system prompts with citation support.

## Frequently Asked Questions

### What is the optional Phase 1.5 in the LLM Wiki retrieval pipeline?

Phase 1.5 is an **optional vector semantic search** that runs between tokenized search and graph expansion. When enabled, it generates embeddings through an OpenAI-compatible API, performs approximate nearest neighbor (ANN) search, and merges results with Phase 1 output before re-ranking. This sub-phase does not constitute a main phase but enhances recall for semantic queries that keyword search might miss.

### How does the graph expansion phase improve retrieval quality?

**Graph Expansion** (Phase 2) improves quality by traversing relationships beyond keyword matches. It uses the top Phase 1 results as seed nodes and applies a 4-signal relevance model to score neighboring pages during a 2-hop traversal. This captures related concepts and implicit connections within the knowledge graph, providing the LLM with broader contextual awareness than lexical search alone.

### What token budget allocations does LLM Wiki support?

The **Budget Control** phase supports context windows from **4,000 to 1,000,000 tokens**. The default allocation distributes approximately **60%** to wiki content, **20%** to chat history, **5%** to the index, and **15%** to system instructions. Users can configure these percentages through the `setContextBudget` API to prioritize specific content types based on their use case.

### How are sources cited in the assembled context?

During **Context Assembly** (Phase 4), the system assigns sequential numbers to selected pages and instructs the LLM to cite sources using bracketed identifiers like `[1]` or `[2]`. The assembled prompt includes full page content (not excerpts) alongside system files such as [`purpose.md`](https://github.com/nashsu/llm_wiki/blob/main/purpose.md) and [`index.md`](https://github.com/nashsu/llm_wiki/blob/main/index.md), ensuring the model can reference specific documentation while generating responses.