How LLM Wiki Implements Budget Control and Context Window Splitting for Retrieval
LLM Wiki enforces strict context window limits by computing a character-level context budget from the model's maxContextSize and hierarchically splitting markdown source into chunks that never exceed the derived per-page cap.
The open-source repository nashsu/llm_wiki solves the fundamental retrieval problem of fitting relevant documentation within finite LLM context windows. By combining percentage-based budget allocation with a structure-aware chunking algorithm, the system guarantees that embedding requests and generation prompts always remain within safe token limits while preserving semantic coherence across document boundaries.
Computing the Context Budget
The foundation of LLM Wiki's budget control system resides in src/lib/context-budget.ts, which exports the pure function computeContextBudget(maxContextSize). This function translates a model's maximum context window into concrete character-count limits that govern every downstream operation.
Percentage-Based Allocation
When invoked, computeContextBudget returns a ContextBudget object that partitions the available context into reserved pools. If the caller passes 0 or undefined, the system defaults to a 200,000-character window (approximately 50K tokens). The allocation follows strict percentages:
maxCtx: The full context window size in characters.responseReserve: Approximately 15% ofmaxCtx, held empty to ensure the model has room to generate answers.indexBudget: Approximately 5% ofmaxCtx, reserved for the wiki-index summary.pageBudget: Approximately 50% ofmaxCtx, representing the total character pool available for all retrieved page content combined.
Per-Page Cap Logic
To prevent any single document from monopolizing the retrieval budget, the system calculates a maxPageSize using the formula:
maxPageSize = min(pageBudget, max(PER_PAGE_FLOOR, floor(pageBudget * 0.3)))
This guarantees that no individual page can consume the entire pageBudget, while still allowing larger chunks when running on long-context models. The PER_PAGE_FLOOR ensures a minimum viable chunk size even on smaller context windows.
Hierarchical Markdown Chunking
With budgets defined, src/lib/text-chunker.ts implements the splitting strategy through its public API chunkMarkdown(content, userOptions?). This function respects document hierarchy and ensures no emitted chunk exceeds the maxPageSize cap derived from the context budget.
Chunking Options and Defaults
The chunker accepts a ChunkingOptions object that controls granularity. The default configuration balances semantic coherence with size constraints:
const DEFAULT_OPTIONS: ChunkingOptions = {
targetChars: 1000, // ideal chunk size
maxChars: 1500, // hard ceiling; oversize chunks are still emitted but flagged
minChars: 200, // tiny chunks are merged into their neighbour
overlapChars: 200, // overlap between adjacent chunks in the same section
};
Structure-Aware Splitting Strategy
The splitting algorithm operates as a priority ladder that respects markdown semantics:
- Heading boundaries (
##,###, etc.) serve as primary split points to isolate sections. - Within sections, the system attempts splits at paragraph boundaries first.
- If paragraphs exceed limits, it falls back to line boundaries, then sentence boundaries, then word boundaries (spaces).
- As a last resort, a hard slice enforces the absolute character limit.
After sizing, the system injects overlapChars (default 200) at chunk boundaries to retain semantic context that spans adjacent sections.
Handling Indivisible Elements
Certain markdown constructs are treated as atomic and never broken:
- Fenced code blocks (triple backtick regions)
- Tables
If these elements exceed maxChars, they are emitted as single oversized chunks with the oversized flag set to true, ensuring the integrity of structured data even when it breaches typical size guidelines.
Integration in the Ingestion Pipeline
The orchestration layer in src/lib/ingest.ts binds budget calculation to chunking. When a source file is ingested, the pipeline executes three sequential steps:
- Invokes
computeContextBudgetto obtain currentpageBudgetandmaxPageSizevalues. - Feeds raw markdown to
chunkMarkdownwithmaxCharsexplicitly set to the computedmaxPageSize. - Dispatches each resulting chunk to the embedding model or vector store.
Because the chunker guarantees that no individual chunk exceeds the per-page cap, the total characters sent to the LLM for any retrieval request never exceed pageBudget, ensuring strict adherence to the overall context window.
Provider-Specific Budget Mapping
In src/lib/llm-providers.ts, the system maps computed budgets to provider-specific request parameters. For models supporting explicit reasoning budgets (such as Anthropic's Claude), the payload includes a thinking.budget_tokens field that mirrors the pageBudget value.
The configuration logic distinguishes between:
max_tokens: Controls visible output length.thinkingBudget: Governs internal chain-of-thought tokens for reasoning models.
This dual-track approach ensures that both the retrieval context and the model's reasoning process respect the originally computed constraints.
Store-Level Enforcement
User-configurable overrides are managed in src/stores/wiki-store.ts, which exposes a budgetTokens property reflecting the calculated budget. The UI presents two modes:
- Auto: Uses the computed
budgetTokensfromcomputeContextBudget. - Custom: Allows users to override the value before the request dispatches.
This layer acts as the final gatekeeper, propagating the definitive budget value to both the chunking subsystem and the LLM provider configuration.
Summary
- Budget Calculation:
src/lib/context-budget.tsderivespageBudget(~50% of context) andmaxPageSize(~30% of page budget) from the model'smaxContextSize, reserving 15% for responses and 5% for the index. - Smart Chunking:
src/lib/text-chunker.tssplits markdown hierarchically at headings, paragraphs, lines, and sentences, never breaking code blocks or tables, and adds 200-character overlaps for context preservation. - Pipeline Integration:
src/lib/ingest.tscoordinates budget computation with chunking to ensure retrieved content fits within the allocated window. - Provider Mapping:
src/lib/llm-providers.tstranslates character budgets into provider-specific token limits, including separate reasoning budgets for supported models. - User Control:
src/stores/wiki-store.tsallows runtime override of computed budgets while maintaining safe defaults.
Frequently Asked Questions
How does LLM Wiki prevent a single large document from consuming the entire context window?
The system enforces a maxPageSize limit calculated as a fraction of the total pageBudget. According to the implementation in src/lib/context-budget.ts, no single page can exceed 30% of the available page budget (subject to a floor value), ensuring that retrieval always leaves room for multiple sources and the model's response reserve.
What happens if a code block or table is larger than the maximum chunk size?
Indivisible constructs are never split. As implemented in src/lib/text-chunker.ts, fenced code blocks and tables are emitted as single chunks even when they exceed maxChars, with the oversized property set to true. This preserves syntax integrity at the cost of temporarily exceeding the target size for that specific element.
Can users override the automatic budget calculations?
Yes. While computeContextBudget provides default allocations, src/stores/wiki-store.ts exposes manual controls through the budgetTokens property. Users can switch from "auto" to "custom" mode to specify exact token limits, which the system then propagates to both the chunking layer and the LLM provider configuration.
Why does the chunker add overlap between adjacent chunks?
The overlapChars parameter (defaulting to 200 characters) ensures that semantic context crossing chunk boundaries is retained in both adjacent pieces. This prevents scenarios where a sentence at the end of one chunk and a sentence at the start of the next lose their contextual relationship, improving retrieval accuracy for questions that span section boundaries.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →