Chunk Size and Overlap Settings for QMD Embeddings: A Technical Deep Dive
QMD determines chunk size and overlap settings based on the embedding model's 900-token limit, a 15% overlap ratio for context continuity, and intelligent boundary detection that prevents splitting code fences or mid-paragraph.
The tobi/qmd repository implements a sophisticated document chunking strategy that balances embedding model constraints with semantic coherence. Understanding the factors that determine chunk size and overlap settings for QMD embeddings is essential for optimizing retrieval performance and maintaining contextual continuity across document boundaries.
Embedding Model Token Limits
The primary constraint driving chunk size is the transformer-based embedding model's maximum token capacity. QMD defaults to a conservative 900 tokens per chunk to ensure compatibility with models like embeddinggemma while leaving safety margin for tokenization variations.
This constant is defined in src/store.ts:
// store.ts, lines 53-55
export const CHUNK_SIZE_TOKENS = 900;
Overlap Calculation and Context Continuity
To prevent semantic discontinuity at chunk boundaries, QMD implements a 15% overlap between consecutive chunks. This ensures that phrases spanning boundaries remain represented in at least one complete chunk, significantly improving recall for semantic search queries.
The overlap calculation uses Math.floor to ensure integer token counts:
// store.ts, line 56
export const CHUNK_OVERLAP_TOKENS = Math.floor(CHUNK_SIZE_TOKENS * 0.15); // ≈135 tokens
Character-Based Approximations
For synchronous processing paths, QMD converts token limits to character limits using an empirical 4:1 character-to-token ratio. This approximation works reliably for mixed prose and code content.
// store.ts, lines 58-59
export const CHUNK_SIZE_CHARS = CHUNK_SIZE_TOKENS * 4; // 3600 characters
export const CHUNK_OVERLAP_CHARS = CHUNK_OVERLAP_TOKENS * 4; // 540 characters
Smart Boundary Detection
QMD employs intelligent break-point detection to ensure chunks end at natural semantic boundaries rather than arbitrary character counts.
Break-Point Scoring and Decay
The findBestCutoff function (lines 84-88 in src/store.ts) evaluates potential break points using regex patterns for headings, paragraphs, lists, and newlines. Each pattern receives a score multiplied by a distance-decay factor of 0.7, penalizing cuts far from the target boundary while favoring nearby semantic breaks.
Code Fence Safety
Chunks must never split inside fenced code blocks (```). The algorithm first detects all fence regions using findCodeFences and isInsideCodeFence (lines 39-68), then discards any break point falling within these protected regions.
Search Window Configuration
To balance granularity with performance, QMD searches for optimal break points within a 200-token window (approximately 800 characters) preceding the target boundary:
// store.ts, lines 60-62
export const CHUNK_WINDOW_TOKENS = 200;
export const CHUNK_WINDOW_CHARS = CHUNK_WINDOW_TOKENS * 4;
Token-Accurate Fallback Mechanism
For absolute compliance with model token limits, chunkDocumentByTokens (lines 124-166) implements a verification step. After creating character-based chunks, it re-tokenizes each chunk. If any chunk exceeds 900 tokens, it recalculates using the actual observed token-to-character ratio for that specific content and resplits accordingly.
Practical Implementation Examples
import { chunkDocument, chunkDocumentByTokens } from "./store.js";
/* 1️⃣ Simple character-based chunking (3600 chars, 540 char overlap) */
const text = await Deno.readTextFile("README.md");
const charChunks = chunkDocument(text);
// → [{ text: "...", pos: 0 }, { text: "...", pos: 3060 }, …]
/* 2️⃣ Token-accurate chunking (900 tokens, 135-token overlap) */
const tokenChunks = await chunkDocumentByTokens(text);
// → [{ text: "...", pos: 0, tokens: 899 }, { text: "...", pos: 2600, tokens: 901 }, …]
/* 3️⃣ Customizing limits for larger models */
import {
CHUNK_SIZE_TOKENS,
CHUNK_OVERLAP_TOKENS,
} from "./store.js";
const largerChunks = chunkDocument(
text,
(CHUNK_SIZE_TOKENS * 4) * 2, // 7200 characters
CHUNK_OVERLAP_TOKENS * 4 * 2 // 1080 characters overlap
);
Summary
- Token limits drive the base chunk size (900 tokens) to ensure embedding model compatibility
- 15% overlap (135 tokens) maintains semantic continuity across chunk boundaries
- 4:1 character-to-token ratio enables fast synchronous processing while approximating token counts
- Smart boundary detection uses a 200-token search window with decay-weighted scoring to find natural break points
- Code fence protection prevents splits inside markdown code blocks using
findCodeFences - Token-accurate fallback re-tokenizes chunks to guarantee strict compliance with the 900-token limit
Frequently Asked Questions
What is the default chunk size in QMD?
QMD defaults to 900 tokens per chunk, defined as CHUNK_SIZE_TOKENS in src/store.ts. This conservative limit ensures compatibility with transformer-based embedding models while leaving margin for tokenization variations. For character-based processing, this translates to approximately 3600 characters using a 4:1 ratio.
How does QMD prevent losing context between chunks?
QMD implements a 15% overlap between consecutive chunks, calculated as Math.floor(CHUNK_SIZE_TOKENS * 0.15) (approximately 135 tokens). This overlap ensures that phrases or concepts spanning chunk boundaries remain fully represented in at least one chunk, significantly improving recall for semantic search queries.
Why does QMD use character-based chunking if embeddings require tokens?
QMD uses character-based chunking (chunkDocument) for synchronous performance during initial processing, applying an empirical 4:1 character-to-token ratio. However, for strict compliance, chunkDocumentByTokens re-tokenizes each chunk and resplits any that exceed the 900-token limit using the actual observed token density, ensuring accuracy without sacrificing speed for the common case.
How does QMD handle code blocks when chunking documents?
QMD protects fenced code blocks (```) from being split across chunks. The algorithm first scans the entire document using findCodeFences to identify all fenced regions, then uses isInsideCodeFence to discard any potential break points that fall within these protected areas. This ensures code examples remain intact and syntactically valid within single chunks.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →