# How LLM Wiki Implements Budget Control and Context Window Splitting for Retrieval

> Discover how LLM Wiki manages retrieval with budget control and context window splitting. Learn about character-level budgeting and hierarchical chunking for efficient context management.

- Repository: [nash_su/llm_wiki](https://github.com/nashsu/llm_wiki)
- Tags: internals
- Published: 2026-09-12

---

**LLM Wiki enforces strict context window limits by computing a character-level context budget from the model's `maxContextSize` and hierarchically splitting markdown source into chunks that never exceed the derived per-page cap.**

The open-source repository `nashsu/llm_wiki` solves the fundamental retrieval problem of fitting relevant documentation within finite LLM context windows. By combining percentage-based budget allocation with a structure-aware chunking algorithm, the system guarantees that embedding requests and generation prompts always remain within safe token limits while preserving semantic coherence across document boundaries.

## Computing the Context Budget

The foundation of LLM Wiki's budget control system resides in [`src/lib/context-budget.ts`](https://github.com/nashsu/llm_wiki/blob/main/src/lib/context-budget.ts), which exports the pure function `computeContextBudget(maxContextSize)`. This function translates a model's maximum context window into concrete character-count limits that govern every downstream operation.

### Percentage-Based Allocation

When invoked, `computeContextBudget` returns a `ContextBudget` object that partitions the available context into reserved pools. If the caller passes `0` or `undefined`, the system defaults to a 200,000-character window (approximately 50K tokens). The allocation follows strict percentages:

- **`maxCtx`**: The full context window size in characters.
- **`responseReserve`**: Approximately 15% of `maxCtx`, held empty to ensure the model has room to generate answers.
- **`indexBudget`**: Approximately 5% of `maxCtx`, reserved for the wiki-index summary.
- **`pageBudget`**: Approximately 50% of `maxCtx`, representing the total character pool available for all retrieved page content combined.

### Per-Page Cap Logic

To prevent any single document from monopolizing the retrieval budget, the system calculates a `maxPageSize` using the formula:

```typescript
maxPageSize = min(pageBudget, max(PER_PAGE_FLOOR, floor(pageBudget * 0.3)))

```

This guarantees that no individual page can consume the entire `pageBudget`, while still allowing larger chunks when running on long-context models. The `PER_PAGE_FLOOR` ensures a minimum viable chunk size even on smaller context windows.

## Hierarchical Markdown Chunking

With budgets defined, [`src/lib/text-chunker.ts`](https://github.com/nashsu/llm_wiki/blob/main/src/lib/text-chunker.ts) implements the splitting strategy through its public API `chunkMarkdown(content, userOptions?)`. This function respects document hierarchy and ensures no emitted chunk exceeds the `maxPageSize` cap derived from the context budget.

### Chunking Options and Defaults

The chunker accepts a `ChunkingOptions` object that controls granularity. The default configuration balances semantic coherence with size constraints:

```typescript
const DEFAULT_OPTIONS: ChunkingOptions = {
  targetChars: 1000,   // ideal chunk size
  maxChars: 1500,      // hard ceiling; oversize chunks are still emitted but flagged
  minChars: 200,       // tiny chunks are merged into their neighbour
  overlapChars: 200,   // overlap between adjacent chunks in the same section
};

```

### Structure-Aware Splitting Strategy

The splitting algorithm operates as a priority ladder that respects markdown semantics:

1. **Heading boundaries** (`##`, `###`, etc.) serve as primary split points to isolate sections.
2. **Within sections**, the system attempts splits at paragraph boundaries first.
3. If paragraphs exceed limits, it falls back to line boundaries, then sentence boundaries, then word boundaries (spaces).
4. As a last resort, a **hard slice** enforces the absolute character limit.

After sizing, the system injects `overlapChars` (default 200) at chunk boundaries to retain semantic context that spans adjacent sections.

### Handling Indivisible Elements

Certain markdown constructs are treated as **atomic** and never broken:

- Fenced code blocks (triple backtick regions)
- Tables

If these elements exceed `maxChars`, they are emitted as single oversized chunks with the `oversized` flag set to `true`, ensuring the integrity of structured data even when it breaches typical size guidelines.

## Integration in the Ingestion Pipeline

The orchestration layer in [`src/lib/ingest.ts`](https://github.com/nashsu/llm_wiki/blob/main/src/lib/ingest.ts) binds budget calculation to chunking. When a source file is ingested, the pipeline executes three sequential steps:

1. Invokes `computeContextBudget` to obtain current `pageBudget` and `maxPageSize` values.
2. Feeds raw markdown to `chunkMarkdown` with `maxChars` explicitly set to the computed `maxPageSize`.
3. Dispatches each resulting chunk to the embedding model or vector store.

Because the chunker guarantees that no individual chunk exceeds the per-page cap, the **total characters sent to the LLM for any retrieval request never exceed `pageBudget`**, ensuring strict adherence to the overall context window.

## Provider-Specific Budget Mapping

In [`src/lib/llm-providers.ts`](https://github.com/nashsu/llm_wiki/blob/main/src/lib/llm-providers.ts), the system maps computed budgets to provider-specific request parameters. For models supporting explicit reasoning budgets (such as Anthropic's Claude), the payload includes a `thinking.budget_tokens` field that mirrors the `pageBudget` value.

The configuration logic distinguishes between:
- **`max_tokens`**: Controls visible output length.
- **`thinkingBudget`**: Governs internal chain-of-thought tokens for reasoning models.

This dual-track approach ensures that both the retrieval context and the model's reasoning process respect the originally computed constraints.

## Store-Level Enforcement

User-configurable overrides are managed in [`src/stores/wiki-store.ts`](https://github.com/nashsu/llm_wiki/blob/main/src/stores/wiki-store.ts), which exposes a `budgetTokens` property reflecting the calculated budget. The UI presents two modes:
- **Auto**: Uses the computed `budgetTokens` from `computeContextBudget`.
- **Custom**: Allows users to override the value before the request dispatches.

This layer acts as the final gatekeeper, propagating the definitive budget value to both the chunking subsystem and the LLM provider configuration.

## Summary

- **Budget Calculation**: [`src/lib/context-budget.ts`](https://github.com/nashsu/llm_wiki/blob/main/src/lib/context-budget.ts) derives `pageBudget` (~50% of context) and `maxPageSize` (~30% of page budget) from the model's `maxContextSize`, reserving 15% for responses and 5% for the index.
- **Smart Chunking**: [`src/lib/text-chunker.ts`](https://github.com/nashsu/llm_wiki/blob/main/src/lib/text-chunker.ts) splits markdown hierarchically at headings, paragraphs, lines, and sentences, never breaking code blocks or tables, and adds 200-character overlaps for context preservation.
- **Pipeline Integration**: [`src/lib/ingest.ts`](https://github.com/nashsu/llm_wiki/blob/main/src/lib/ingest.ts) coordinates budget computation with chunking to ensure retrieved content fits within the allocated window.
- **Provider Mapping**: [`src/lib/llm-providers.ts`](https://github.com/nashsu/llm_wiki/blob/main/src/lib/llm-providers.ts) translates character budgets into provider-specific token limits, including separate reasoning budgets for supported models.
- **User Control**: [`src/stores/wiki-store.ts`](https://github.com/nashsu/llm_wiki/blob/main/src/stores/wiki-store.ts) allows runtime override of computed budgets while maintaining safe defaults.

## Frequently Asked Questions

### How does LLM Wiki prevent a single large document from consuming the entire context window?

The system enforces a `maxPageSize` limit calculated as a fraction of the total `pageBudget`. According to the implementation in [`src/lib/context-budget.ts`](https://github.com/nashsu/llm_wiki/blob/main/src/lib/context-budget.ts), no single page can exceed 30% of the available page budget (subject to a floor value), ensuring that retrieval always leaves room for multiple sources and the model's response reserve.

### What happens if a code block or table is larger than the maximum chunk size?

Indivisible constructs are never split. As implemented in [`src/lib/text-chunker.ts`](https://github.com/nashsu/llm_wiki/blob/main/src/lib/text-chunker.ts), fenced code blocks and tables are emitted as single chunks even when they exceed `maxChars`, with the `oversized` property set to `true`. This preserves syntax integrity at the cost of temporarily exceeding the target size for that specific element.

### Can users override the automatic budget calculations?

Yes. While `computeContextBudget` provides default allocations, [`src/stores/wiki-store.ts`](https://github.com/nashsu/llm_wiki/blob/main/src/stores/wiki-store.ts) exposes manual controls through the `budgetTokens` property. Users can switch from "auto" to "custom" mode to specify exact token limits, which the system then propagates to both the chunking layer and the LLM provider configuration.

### Why does the chunker add overlap between adjacent chunks?

The `overlapChars` parameter (defaulting to 200 characters) ensures that semantic context crossing chunk boundaries is retained in both adjacent pieces. This prevents scenarios where a sentence at the end of one chunk and a sentence at the start of the next lose their contextual relationship, improving retrieval accuracy for questions that span section boundaries.