How QMD's Smart Chunking Algorithm Identifies Markdown Boundaries for Optimal Tokenization

QMD's smart chunking algorithm scans markdown documents for structural break points, scores them by semantic importance, and applies a distance-decay function to select the optimal cut position within a configurable window, ensuring chunks respect natural document boundaries while maximizing token efficiency.

QMD (Query-Markup-Documents) is an open-source semantic search tool that processes large markdown files for embedding and reranking. The smart chunking algorithm implemented in tobi/qmd ensures that document splits occur at meaningful structural boundaries rather than arbitrary character limits, preserving semantic coherence for vector search applications.

How the Smart Chunking Pipeline Works

The algorithm operates through three distinct stages to transform raw markdown into optimally-sized semantic chunks.

First, the system pre-scans the entire document to identify potential break points including headings, horizontal rules, blank lines, list items, and code-fence boundaries. Next, it scores each break point using a fixed hierarchy where h1 headings receive the highest priority (score 100), followed by h2 (90), descending through h6, code delimiters, horizontal rules, blank lines, and list items. Finally, the algorithm selects the best cut position within a configurable window (approximately 800 characters or 200 tokens) using a distance-decay function that favors high-scoring points while preferring earlier positions when scores are similar.

Break Point Detection and Scoring

The foundation of QMD's boundary detection lies in the BREAK_PATTERNS constant defined in src/store.ts at lines 92-104.

This ordered array pairs regular expressions with numeric scores and type labels. Headings receive premium scores (h1 = 100, h2 = 90, h3 = 80, h4 = 70, h5 = 60, h6 = 50), ensuring section boundaries take precedence over structural elements like code-block delimiters (40), horizontal rules (30), blank lines (20), and list items (10). A minimal newline fallback (score 1) guarantees the algorithm always finds a cut point even in dense text.

The scanBreakPoints(text) function (lines 112-133) iterates over every pattern, recording character offsets via match.index. When multiple patterns match the same position, the system retains only the highest-scoring entry, ensuring a line that matches both a heading pattern and a newline pattern is classified as a heading.

Code Fence Protection

To prevent syntax errors in embedded code snippets, QMD explicitly excludes break points that fall within code blocks.

The findCodeFences(text) function (lines 139-155) extracts all regions bounded by triple backticks (```), recording their start and end positions. During the cut selection phase, the isInsideCodeFence check filters out any break points that fall within these fenced regions, ensuring the algorithm never splits a code block in half.

The Distance-Decay Selection Algorithm

The core optimization logic resides in findBestCutoff (lines 170-190), which implements a quadratic distance-decay function to balance boundary quality against token efficiency.

Given a target end position (targetCharPos), a search window (windowChars ≈ 800), and a decay factor, the function examines break points backward from the target. For each candidate, it calculates:

const normalizedDist = (targetCharPos - bp.pos) / windowChars;
const multiplier = 1.0 - (normalizedDist * normalizedDist) * decayFactor;
const finalScore = bp.score * multiplier;

The quadratic decay (normalized distance squared) ensures that high-scoring boundaries slightly before the target strongly outrank low-scoring boundaries at the exact limit. For example, an h2 heading (score 90) at 85% of the window distance retains a higher final score than a blank line (score 20) at 95% distance, even though the blank line is closer to the target.

Token-Accurate Chunking

For applications requiring strict token limits (such as embedding models with fixed context windows), QMD provides chunkDocumentByTokens (lines 223-267).

This function implements a two-pass verification strategy:

  1. Initial estimation: Run the character-based algorithm using a conservative estimate of 3 characters per token to produce candidate chunks.
  2. Token verification: Tokenize each chunk using the Llama-CPP tokenizer. If a chunk exceeds the limit, recalculate using the actual chars-per-token ratio derived from the overshoot, then re-split and retokenize.

This hybrid approach balances speed (avoiding tokenization of the entire document upfront) with accuracy (guaranteeing no chunk exceeds the model's limit).

Practical Implementation Examples

Basic Character-Based Chunking

The default chunking method optimizes for semantic boundaries without strict token guarantees:

import { chunkDocument } from "qmd/src/store";

const markdown = await Deno.readTextFile("documentation.md");
const chunks = chunkDocument(markdown, 3600, 540, 800);

// Returns: [{ text: "...", pos: 0 }, { text: "...", pos: 3060 }, ...]
console.log(`Generated ${chunks.length} semantic chunks`);

Parameters specify maximum characters (3600), overlap characters (540 for 15% overlap), and search window (800 characters).

Token-Accurate Processing for Embeddings

For LLM pipelines requiring strict token limits:

import { chunkDocumentByTokens } from "qmd/src/store";

async function prepareForEmbedding(filePath: string) {
  const content = await Deno.readTextFile(filePath);
  const chunks = await chunkDocumentByTokens(content, 900, 135);
  
  // Each chunk guaranteed ≤ 900 tokens
  return chunks.map(c => ({ 
    text: c.text, 
    tokenCount: c.tokens 
  }));
}

CLI Usage

The qmd command-line tool exposes these algorithms through simple flags:


# Default smart chunking (character-based with markdown awareness)

qmd collection add docs --name knowledge-base

# Strict token limits for embedding models

qmd embed --tokenize --max-tokens 512

Summary

QMD's smart chunking algorithm delivers semantic document segmentation through these key mechanisms:

  • Hierarchical boundary scoring assigns priority to markdown structural elements (headings > code fences > horizontal rules > blank lines) to ensure cuts occur at meaningful semantic boundaries.
  • Quadratic distance-decay optimization balances the desire for high-quality boundaries against token efficiency, preferring a high-scoring boundary slightly before the limit over a low-scoring boundary at the exact character count.
  • Code fence protection explicitly excludes break points within triple-backtick regions to prevent syntax errors in embedded code snippets.
  • Two-pass token verification in chunkDocumentByTokens combines fast character-based estimation with accurate LLM tokenizer validation to guarantee strict token limits without sacrificing boundary quality.

Frequently Asked Questions

How does QMD prevent splitting code blocks during chunking?

QMD uses the findCodeFences function in src/store.ts (lines 139-155) to pre-identify all regions bounded by triple backticks. During the findBestCutoff selection process, any break point that falls within these fenced regions is filtered out via the isInsideCodeFence check, ensuring code blocks remain intact across chunk boundaries.

What is the distance-decay function and why is it quadratic?

The distance-decay function calculates a score multiplier based on how far a break point is from the target chunk end position. QMD uses a quadratic formula (normalizedDist * normalizedDist) rather than linear decay because it more aggressively penalizes boundaries near the window edge while preserving high scores for quality boundaries slightly before the target. This ensures an h2 heading at 85% of the window distance outranks a blank line at 95%, even though the blank line is closer to the limit.

Can I adjust the chunk size and overlap parameters?

Yes, both chunkDocument and chunkDocumentByTokens accept configurable parameters. For character-based chunking, you can specify maxChars (default ~3600), overlapChars (default 15% of max), and windowChars (default ~800). For token-based chunking, you provide maxTokens (default 900) and overlapTokens (default 135). These parameters are defined in src/store.ts lines 156-176 and 223-267 respectively.

How does the token-accurate chunking differ from character-based chunking?

Character-based chunking (chunkDocument) uses a conservative estimate of 3 characters per token and optimizes for markdown boundaries without strict token guarantees. Token-accurate chunking (chunkDocumentByTokens) implements a two-pass strategy: it first runs the character algorithm, then tokenizes each chunk using the Llama-CPP tokenizer. Any chunk exceeding the limit is re-split using the actual character-to-token ratio and retokenized, guaranteeing strict compliance with model context windows while preserving semantic boundary quality.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →