# How QMD's Smart Chunking Preserves Markdown Structure: A Technical Deep Dive

> Discover how QMD's smart chunking preserves markdown structure. Learn the technical details of boundary scoring, fence protection, and decay algorithms for optimal code splitting that respects your formatting.

- Repository: [Tobias Lütke/qmd](https://github.com/tobi/qmd)
- Tags: deep-dive
- Published: 2026-02-16

---

**QMD's smart chunking algorithm preserves markdown structure by scoring boundary patterns, protecting code fences, and applying distance-weighted decay to select optimal split points, ensuring chunks respect headings and formatting while maintaining size limits.**

QMD (Query Markdown) is an open-source tool by Tobias Lütke (`tobi/qmd`) that indexes and searches markdown documents using vector embeddings. At its core, **smart chunking** solves the critical problem of splitting large documents into embeddable pieces without breaking logical structure, ensuring that headings, lists, and code blocks remain intact within each chunk for better semantic search results.

## The Five-Stage Chunking Pipeline

Unlike naive character-count splitting that can sever words or break code blocks, QMD implements a multi-stage pipeline in [`src/store.ts`](https://github.com/tobi/qmd/blob/main/src/store.ts) that understands markdown syntax. The algorithm evaluates potential break points using regex patterns, excludes invalid locations inside code fences, and applies mathematical scoring to find the most semantically appropriate split location near the target chunk size.

### 1. Detecting Markdown Break Points with BREAK_PATTERNS

The algorithm first scans the entire document using `BREAK_PATTERNS`—a set of regular expressions that identify natural markdown boundaries. These patterns detect headings, horizontal rules, blank lines, list items, and the start of code fences. Each match generates a `BreakPoint` object assigned a **score** reflecting its suitability as a split location, with headings receiving the highest priority.

*Source:* [`src/store.ts`](https://github.com/tobi/qmd/blob/main/src/store.ts) lines 92-104

### 2. Protecting Code Fence Integrity

To prevent splitting inside code blocks, the `findCodeFences` function (lines 138-160) walks the text and records every region between triple backticks (``` … ```). The subsequent split algorithm uses `isInsideCodeFence` to discard any break point falling within these regions, ensuring code examples remain complete and syntactically valid within a single chunk.

### 3. Selecting Optimal Cut Locations with Distance Decay

For each target chunk size, `findBestCutoff` (lines 183-216) looks backward through a configurable window (approximately 200 tokens) and evaluates candidate break points. The algorithm applies a **distance-decay factor** using a squared distance term, allowing break points farther from the target to retain decent weight while favoring locations nearer the ideal size. The highest-scoring valid position becomes the chunk boundary.

### 4. Iterating with Context Overlap

The `chunkDocument` function (lines 57-118) repeatedly slices text from `charPos` to the chosen `endPos`. After each slice, it **overlaps** the next chunk by approximately 15% of the chunk size (configurable) to preserve context at boundaries. This overlap ensures that semantic meaning isn't lost where chunks meet, maintaining coherence for embedding models.

### 5. Token-Accurate Fallback Processing

When strict token limits are required, `chunkDocumentByTokens` (lines 1240-1256) first runs the character-based algorithm with a generous estimate (approximately 4 characters per token). Any chunk exceeding the token budget is re-split using the actual token count, preserving the same markdown-aware breakpoint logic but with tighter constraints.

## Implementation Details in src/store.ts

The core logic resides in [`src/store.ts`](https://github.com/tobi/qmd/blob/main/src/store.ts), which exports the public APIs `chunkDocument` and `chunkDocumentByTokens`. The scoring system prioritizes structural elements—headings score highest, followed by list items and horizontal rules—while the squared distance decay ensures the algorithm doesn't force splits at poor locations just to hit exact byte counts. This mathematical approach balances structural preservation against size constraints.

## Practical Usage Examples

You can use QMD's chunking functions directly in your own TypeScript applications:

```typescript
import { chunkDocument, chunkDocumentByTokens } from "./store";

// Character-based smart chunking (preserves markdown structure)
const markdown = await Deno.readTextFile("README.md");
const charChunks = chunkDocument(markdown);
// each element: { text: string, pos: number }
console.log(charChunks.map(c => c.text.length));

// Token-accurate chunking (still uses the same markdown-aware logic)
const tokenChunks = await chunkDocumentByTokens(markdown, 900, 135, 200);
console.log(tokenChunks.map(c => c.tokens));

```

## Integration with the QMD Pipeline

According to the source code in [`src/qmd.ts`](https://github.com/tobi/qmd/blob/main/src/qmd.ts), the `chunkDocumentByTokens` function is called during the indexing phase before documents are embedded and stored. This integration ensures that the vector database receives semantically coherent chunks rather than arbitrary text segments, improving retrieval accuracy when querying the markdown corpus.

## Summary

- **Pattern-based detection** uses `BREAK_PATTERNS` to score markdown boundaries, prioritizing headings and structural elements.
- **Code fence protection** via `findCodeFences` and `isInsideCodeFence` ensures code blocks are never split mid-content.
- **Distance-weighted scoring** in `findBestCutoff` uses squared decay to balance proximity to target size against boundary quality.
- **Context overlap** of approximately 15% between chunks preserves semantic continuity at boundaries.
- **Token fallback** in `chunkDocumentByTokens` provides exact token limits when needed while retaining markdown awareness.

## Frequently Asked Questions

### What makes QMD's chunking "smart" compared to simple character splitting?

QMD's algorithm understands markdown syntax through regex patterns that identify headings, lists, and code fences, whereas simple splitting severs text at fixed character counts regardless of content structure. This semantic awareness ensures that `chunkDocument` never breaks a code block or splits a heading from its content, producing self-contained readable fragments.

### How does QMD prevent splitting inside code blocks?

The `findCodeFences` function pre-scans the document to record all triple-backtick regions (lines 138-160 in [`src/store.ts`](https://github.com/tobi/qmd/blob/main/src/store.ts)). During the split selection phase, any break point falling inside these recorded regions is automatically discarded by `isInsideCodeFence`, guaranteeing that code examples remain syntactically complete within a single chunk.

### Why does the algorithm use a squared distance decay factor?

The squared distance term in `findBestCutoff` (lines 183-216) ensures that break points farther from the target chunk size retain meaningful weight in the scoring calculation while still favoring closer locations. This prevents the algorithm from selecting poor boundaries simply because they hit an exact byte count, allowing flexibility to find the nearest high-quality markdown boundary.

### Can I adjust the chunk size and overlap parameters?

Yes, both `chunkDocument` and `chunkDocumentByTokens` accept configurable parameters for target size and overlap. The default overlap is approximately 15% of the chunk size, and the search window for break points defaults to roughly 200 tokens, but these can be tuned based on your specific embedding model's context window and retrieval requirements.