How QMD Detects Document Changes and Handles Updates: A Complete Technical Guide
QMD detects document changes by computing SHA-256 hashes of file contents and comparing them against stored values in a SQLite database, handling updates through a four-state logic (unchanged, modified, new, or removed) implemented in src/qmd.ts and src/store.ts.
When managing large document collections, efficiently detecting changes is critical for maintaining an accurate search index. The tobi/qmd project implements a robust content-addressable storage system that uses cryptographic hashing to detect document changes and SQLite transactions to handle updates atomically. This article explores the complete technical implementation of how QMD detects document modifications and synchronizes its internal state.
The Core Change Detection Mechanism
QMD's change detection relies on a deterministic content hashing strategy combined with a relational database lookup. When you run qmd update, the system walks the configured collection paths and processes each file through a standardized pipeline.
Computing Content Hashes with SHA-256
At the heart of the detection logic is the hashContent() function in src/store.ts. This utility generates a canonical identifier for any file by computing a SHA-256 digest over the raw byte content:
// src/store.ts – hashContent()
export async function hashContent(content: string): Promise<string> {
const hash = createHash("sha256");
hash.update(content);
return hash.digest("hex");
}
The resulting 64-character hexadecimal string serves as the content hash. Because the hash is computed over raw contents without line-ending conversion, even subtle changes (like adding a single character) produce a completely different hash, ensuring precise change detection.
The SQLite Index Lookup
Once the hash is computed, QMD queries the SQLite database to determine the file's current status. The findActiveDocument() function in src/store.ts retrieves the stored record for any active document:
// src/store.ts – findActiveDocument()
export function findActiveDocument(
db: Database,
collectionName: string,
path: string
): { id: number; hash: string; title: string } | null {
const row = db.prepare(`
SELECT id, hash, title FROM documents
WHERE collection = ? AND path = ? AND active = 1
`).get(collectionName, path) as { id: number; hash: string; title: string } | undefined;
return row ?? null;
}
This lookup returns the document's internal ID, the previously stored hash, and the cached title. If no record exists, the function returns null, signaling that this is a new file.
The Four-State Update Logic
Based on the hash comparison and database lookup, QMD categorizes each file into one of four states and executes the corresponding database operations. This logic is implemented in the indexFiles() function within src/qmd.ts.
Unchanged Files (Hash Match)
When the computed hash matches the stored hash, the file content is identical to the indexed version. However, QMD still checks if the title (extracted from frontmatter or headings) has changed:
// src/qmd.ts – inside the main for-loop
if (existing.hash === hash) {
// ---------- unchanged ----------
if (existing.title !== title) {
// only title changed → update title
updateDocumentTitle(db, existing.id, title, now);
updated++;
} else {
unchanged++;
}
}
If only the title differs, QMD calls updateDocumentTitle() to refresh the metadata without touching the content hash, minimizing database writes.
Modified Files (Hash Mismatch)
When the hashes differ, QMD treats the file as modified. It inserts the new content into the content table (content-addressable storage) and updates the document record with the new hash, title, and modification timestamp:
} else {
// ---------- content changed ----------
insertContent(db, hash, content, now);
const stat = statSync(filepath);
updateDocument(db, existing.id, title, hash,
stat ? new Date(stat.mtime).toISOString() : now);
updated++;
}
This approach maintains a complete history of content hashes while keeping the document metadata current.
New Files (No Existing Record)
For files not present in the database, QMD performs an initial insertion. It creates both a content row and a document row, capturing the creation time and modification time from the filesystem:
} else {
// ---------- brand-new file ----------
indexed++;
insertContent(db, hash, content, now);
const stat = statSync(filepath);
insertDocument(db, collectionName, path, title, hash,
stat ? new Date(stat.birthtime).toISOString() : now,
stat ? new Date(stat.mtime).toISOString() : now);
}
Removed Files (Deactivation)
After processing all files found on disk, QMD identifies documents that exist in the database but were not encountered during the filesystem walk. These are marked as inactive (soft-deleted):
// src/qmd.ts – after the walk finishes
const allActive = getActiveDocumentPaths(db, collectionName);
let removed = 0;
for (const path of allActive) {
if (!seenPaths.has(path)) {
deactivateDocument(db, collectionName, path);
removed++;
}
}
The seenPaths set contains every path processed during the collection walk. Any active document whose path is absent from this set is considered deleted and deactivated via deactivateDocument(), which sets the active flag to 0.
Housekeeping and Cleanup Operations
After processing additions, updates, and deletions, QMD performs maintenance to keep the database lean.
Pruning Orphaned Content Hashes
When documents are deactivated or updated to new content hashes, old hash entries in the content table may become unreferenced. The cleanupOrphanedContent() function removes these dangling rows:
// src/store.ts – cleanupOrphanedContent()
export function cleanupOrphanedContent(db: Database): number {
const result = db.prepare(`
DELETE FROM content
WHERE hash NOT IN (SELECT DISTINCT hash FROM documents WHERE active = 1)
`).run();
return result.changes;
}
This SQL statement deletes any content hash not currently referenced by an active document, returning the count of removed rows for the final summary report.
Code Implementation Walkthrough
The following examples demonstrate how to interact with QMD's change detection system programmatically.
Manually Invoking the Update Routine (CLI)
# Update all collections (detects adds/changes/removals)
qmd update
Using the Library Programmatically
import { createStore, getDefaultDbPath } from "./store.js";
// Open the store (default ~/.cache/qmd/index.sqlite)
const store = createStore(); // ← creates DB if needed
const db = store.db;
// Path to a markdown file you want to index
const filePath = "/home/user/notes/todo.md";
// Read the file, compute its hash, etc.
import { readFileSync } from "fs";
import { hashContent, insertContent, findActiveDocument,
updateDocument, insertDocument } from "./store.js";
const raw = readFileSync(filePath, "utf-8");
const hash = await hashContent(raw);
const title = "My TODOs"; // normally extracted with extractTitle()
const collection = "notes";
const relPath = "todo.md";
const now = new Date().toISOString();
const existing = findActiveDocument(db, collection, relPath);
if (existing) {
if (existing.hash !== hash) {
// file changed – update
insertContent(db, hash, raw, now);
updateDocument(db, existing.id, title, hash, now);
} else {
// unchanged – maybe just update the title
// (optional) updateDocumentTitle(db, existing.id, title, now);
}
} else {
// new file
insertContent(db, hash, raw, now);
insertDocument(db, collection, relPath, title, hash, now, now);
}
Detecting a Missing File (Deactivation)
import { getActiveDocumentPaths, deactivateDocument } from "./store.js";
const active = getActiveDocumentPaths(db, "notes");
const seen = new Set(["todo.md", "journal/2023/entry.md"]); // paths encountered during a scan
for (const path of active) {
if (!seen.has(path)) {
deactivateDocument(db, "notes", path); // marks document inactive
}
}
Summary
- SHA-256 hashing provides the foundation for change detection in
src/store.ts, generating unique content identifiers that detect even single-character modifications. - Four-state logic in
src/qmd.tscategorizes files as unchanged, modified, new, or removed, applying targeted database operations for each state to minimize write overhead. - Soft deletion via the
activeflag allows QMD to track removed files without losing historical data, whilecleanupOrphanedContent()maintains database hygiene by pruning unreferenced hash rows. - Atomic SQLite operations ensure data consistency during updates, with separate helper functions in
src/store.tshandling inserts, updates, and deactivations through parameterized queries.
Frequently Asked Questions
How does QMD determine if a file has been modified?
QMD determines file modifications by comparing SHA-256 hashes. When indexFiles() processes a file, it calls hashContent() in src/store.ts to compute a hash of the raw file contents. It then compares this value against the hash stored in the SQLite database via findActiveDocument(). If the hashes differ, QMD treats the file as modified and updates both the content table and document record.
What happens to documents that are deleted from the filesystem?
When a file is deleted from the filesystem, QMD detects this during the cleanup phase of indexFiles() in src/qmd.ts. After processing all existing files, the system retrieves all active document paths using getActiveDocumentPaths() and compares them against a Set of paths seen during the walk. Any active document whose path is not in the seen set is deactivated via deactivateDocument(), which sets the active flag to 0 in the database, effectively performing a soft delete.
Does QMD store multiple versions of document content?
Yes, QMD implements content-addressable storage that preserves distinct content versions. When a file is modified, insertContent() in src/store.ts inserts a new row into the content table with the new SHA-256 hash and file contents, without deleting the old hash entry. However, during the housekeeping phase, cleanupOrphanedContent() removes any content rows that are no longer referenced by active documents, pruning historical versions that belong to deleted or updated files.
How does QMD handle title changes without content modifications?
QMD optimizes for title-only updates by checking the title separately from the content hash. When indexFiles() finds that the computed hash matches the stored hash (indicating unchanged content), it compares the extracted title against the stored title. If they differ, it calls updateDocumentTitle() in src/store.ts to update only the title field in the database, avoiding unnecessary writes to the content table and preserving the existing hash relationship.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →