# How QMD's Smart Chunking Algorithm Identifies Markdown Boundaries for Optimal Tokenization

> Discover how QMD's smart chunking algorithm finds markdown boundaries for efficient tokenization. Learn about its structural break point scanning, semantic scoring, and distance-decay function.

- Repository: [Tobias Lütke/qmd](https://github.com/tobi/qmd)
- Tags: internals
- Published: 2026-02-15

---

**QMD's smart chunking algorithm scans markdown documents for structural break points, scores them by semantic importance, and applies a distance-decay function to select the optimal cut position within a configurable window, ensuring chunks respect natural document boundaries while maximizing token efficiency.**

QMD (Query-Markup-Documents) is an open-source semantic search tool that processes large markdown files for embedding and reranking. The **smart chunking algorithm** implemented in `tobi/qmd` ensures that document splits occur at meaningful structural boundaries rather than arbitrary character limits, preserving semantic coherence for vector search applications.

## How the Smart Chunking Pipeline Works

The algorithm operates through three distinct stages to transform raw markdown into optimally-sized semantic chunks.

First, the system **pre-scans the entire document** to identify potential break points including headings, horizontal rules, blank lines, list items, and code-fence boundaries. Next, it **scores each break point** using a fixed hierarchy where `h1` headings receive the highest priority (score 100), followed by `h2` (90), descending through `h6`, code delimiters, horizontal rules, blank lines, and list items. Finally, the algorithm **selects the best cut position** within a configurable window (approximately 800 characters or 200 tokens) using a distance-decay function that favors high-scoring points while preferring earlier positions when scores are similar.

## Break Point Detection and Scoring

The foundation of QMD's boundary detection lies in the `BREAK_PATTERNS` constant defined in [`src/store.ts`](https://github.com/tobi/qmd/blob/main/src/store.ts) at lines 92-104.

This ordered array pairs regular expressions with numeric scores and type labels. Headings receive premium scores (`h1` = 100, `h2` = 90, `h3` = 80, `h4` = 70, `h5` = 60, `h6` = 50), ensuring section boundaries take precedence over structural elements like code-block delimiters (40), horizontal rules (30), blank lines (20), and list items (10). A minimal newline fallback (score 1) guarantees the algorithm always finds a cut point even in dense text.

The `scanBreakPoints(text)` function (lines 112-133) iterates over every pattern, recording character offsets via `match.index`. When multiple patterns match the same position, the system retains only the **highest-scoring** entry, ensuring a line that matches both a heading pattern and a newline pattern is classified as a heading.

## Code Fence Protection

To prevent syntax errors in embedded code snippets, QMD explicitly excludes break points that fall within code blocks.

The `findCodeFences(text)` function (lines 139-155) extracts all regions bounded by triple backticks (```` ``` ````), recording their start and end positions. During the cut selection phase, the `isInsideCodeFence` check filters out any break points that fall within these fenced regions, ensuring the algorithm never splits a code block in half.

## The Distance-Decay Selection Algorithm

The core optimization logic resides in `findBestCutoff` (lines 170-190), which implements a quadratic distance-decay function to balance boundary quality against token efficiency.

Given a target end position (`targetCharPos`), a search window (`windowChars` ≈ 800), and a decay factor, the function examines break points backward from the target. For each candidate, it calculates:

```ts
const normalizedDist = (targetCharPos - bp.pos) / windowChars;
const multiplier = 1.0 - (normalizedDist * normalizedDist) * decayFactor;
const finalScore = bp.score * multiplier;

```

The **quadratic decay** (normalized distance squared) ensures that high-scoring boundaries slightly before the target strongly outrank low-scoring boundaries at the exact limit. For example, an `h2` heading (score 90) at 85% of the window distance retains a higher final score than a blank line (score 20) at 95% distance, even though the blank line is closer to the target.

## Token-Accurate Chunking

For applications requiring strict token limits (such as embedding models with fixed context windows), QMD provides `chunkDocumentByTokens` (lines 223-267).

This function implements a **two-pass verification strategy**:

1. **Initial estimation**: Run the character-based algorithm using a conservative estimate of 3 characters per token to produce candidate chunks.
2. **Token verification**: Tokenize each chunk using the Llama-CPP tokenizer. If a chunk exceeds the limit, recalculate using the actual `chars-per-token` ratio derived from the overshoot, then re-split and retokenize.

This hybrid approach balances speed (avoiding tokenization of the entire document upfront) with accuracy (guaranteeing no chunk exceeds the model's limit).

## Practical Implementation Examples

### Basic Character-Based Chunking

The default chunking method optimizes for semantic boundaries without strict token guarantees:

```ts
import { chunkDocument } from "qmd/src/store";

const markdown = await Deno.readTextFile("documentation.md");
const chunks = chunkDocument(markdown, 3600, 540, 800);

// Returns: [{ text: "...", pos: 0 }, { text: "...", pos: 3060 }, ...]
console.log(`Generated ${chunks.length} semantic chunks`);

```

Parameters specify maximum characters (3600), overlap characters (540 for 15% overlap), and search window (800 characters).

### Token-Accurate Processing for Embeddings

For LLM pipelines requiring strict token limits:

```ts
import { chunkDocumentByTokens } from "qmd/src/store";

async function prepareForEmbedding(filePath: string) {
  const content = await Deno.readTextFile(filePath);
  const chunks = await chunkDocumentByTokens(content, 900, 135);
  
  // Each chunk guaranteed ≤ 900 tokens
  return chunks.map(c => ({ 
    text: c.text, 
    tokenCount: c.tokens 
  }));
}

```

### CLI Usage

The `qmd` command-line tool exposes these algorithms through simple flags:

```bash

# Default smart chunking (character-based with markdown awareness)

qmd collection add docs --name knowledge-base

# Strict token limits for embedding models

qmd embed --tokenize --max-tokens 512

```

## Summary

QMD's smart chunking algorithm delivers semantic document segmentation through these key mechanisms:

- **Hierarchical boundary scoring** assigns priority to markdown structural elements (headings > code fences > horizontal rules > blank lines) to ensure cuts occur at meaningful semantic boundaries.
- **Quadratic distance-decay optimization** balances the desire for high-quality boundaries against token efficiency, preferring a high-scoring boundary slightly before the limit over a low-scoring boundary at the exact character count.
- **Code fence protection** explicitly excludes break points within triple-backtick regions to prevent syntax errors in embedded code snippets.
- **Two-pass token verification** in `chunkDocumentByTokens` combines fast character-based estimation with accurate LLM tokenizer validation to guarantee strict token limits without sacrificing boundary quality.

## Frequently Asked Questions

### How does QMD prevent splitting code blocks during chunking?

QMD uses the `findCodeFences` function in [`src/store.ts`](https://github.com/tobi/qmd/blob/main/src/store.ts) (lines 139-155) to pre-identify all regions bounded by triple backticks. During the `findBestCutoff` selection process, any break point that falls within these fenced regions is filtered out via the `isInsideCodeFence` check, ensuring code blocks remain intact across chunk boundaries.

### What is the distance-decay function and why is it quadratic?

The distance-decay function calculates a score multiplier based on how far a break point is from the target chunk end position. QMD uses a quadratic formula (`normalizedDist * normalizedDist`) rather than linear decay because it more aggressively penalizes boundaries near the window edge while preserving high scores for quality boundaries slightly before the target. This ensures an `h2` heading at 85% of the window distance outranks a blank line at 95%, even though the blank line is closer to the limit.

### Can I adjust the chunk size and overlap parameters?

Yes, both `chunkDocument` and `chunkDocumentByTokens` accept configurable parameters. For character-based chunking, you can specify `maxChars` (default ~3600), `overlapChars` (default 15% of max), and `windowChars` (default ~800). For token-based chunking, you provide `maxTokens` (default 900) and `overlapTokens` (default 135). These parameters are defined in [`src/store.ts`](https://github.com/tobi/qmd/blob/main/src/store.ts) lines 156-176 and 223-267 respectively.

### How does the token-accurate chunking differ from character-based chunking?

Character-based chunking (`chunkDocument`) uses a conservative estimate of 3 characters per token and optimizes for markdown boundaries without strict token guarantees. Token-accurate chunking (`chunkDocumentByTokens`) implements a two-pass strategy: it first runs the character algorithm, then tokenizes each chunk using the Llama-CPP tokenizer. Any chunk exceeding the limit is re-split using the actual character-to-token ratio and retokenized, guaranteeing strict compliance with model context windows while preserving semantic boundary quality.