How QMD's Docid System Generates 6-Character Content Hashes
QMD generates deterministic 6-character document identifiers by computing the SHA-256 hash of the entire file content and extracting the first six hexadecimal characters, enabling content-addressed storage without relying on filenames or timestamps.
The tobi/qmd repository implements a content-addressed storage system where every document receives a stable, unique identifier derived solely from its text. Unlike traditional systems that rely on file paths or timestamps, QMD's docid system ensures that identical content always produces identical identifiers, while any modification—down to a single character—generates an entirely new hash. This approach centers on two core operations: computing a full SHA-256 digest and truncating it to a human-friendly 6-character prefix.
How QMD's Docid System Works
The generation pipeline follows a strict two-step process implemented in src/store.ts. First, the system computes a cryptographic hash of the entire document body. Second, it extracts a short, printable identifier from that hash for user interaction and database indexing.
Content Hashing with SHA-256
The hashContent function in src/store.ts (lines 1919–1923) creates a standard SHA-256 digest of the raw document string:
export async function hashContent(content: string): Promise<string> {
const hash = createHash("sha256");
hash.update(content);
return hash.digest("hex");
}
This function returns the complete 64-character hexadecimal string representing the document's cryptographic fingerprint. The system stores this full hash in the database to guarantee uniqueness, while the shorter docid serves as a user-facing handle.
Extracting the 6-Character Identifier
Once the full hash is computed, the getDocid function (lines 944–948) truncates it to the first six characters:
export function getDocid(hash: string): string {
return hash.slice(0, 6);
}
This produces values like e3b0c4, which QMD presents to users as #e3b0c4. The 6-character length provides a practical balance between readability and collision resistance for typical document collections.
Database Storage and Retrieval
The docid system integrates tightly with QMD's SQLite storage layer. Rather than storing the short identifier separately, the system derives it on demand from the persisted full hash, ensuring consistency across all queries.
Storing the Full Hash
When indexing a document, QMD inserts the 64-character SHA-256 hash into the content table and references it from the documents table. The short docid is never stored as a distinct column; instead, the system relies on the LIKE operator to match prefixes against the full hash column.
Looking Up Documents by Docid
The findDocumentByDocid function (lines 1535–1548) implements reverse lookup using SQL pattern matching:
export function findDocumentByDocid(db: Database, docid: string) {
const shortHash = normalizeDocid(docid);
return db.prepare(`
SELECT 'qmd://' || d.collection || '/' || d.path AS filepath,
d.hash
FROM documents d
WHERE d.hash LIKE ? AND d.active = 1
LIMIT 1
`).get(`${shortHash}%`);
}
This query appends a wildcard (%) to the 6-character prefix and searches for the first active document with a matching hash start. Collision handling is intentionally simple: if multiple documents share the same 6-character prefix, the query returns the first match.
Normalizing User Input
To accommodate various user input formats, the normalizeDocid helper strips optional decorators. It removes surrounding quotes and leading # characters, allowing users to reference documents as #e3b0c4, "e3b0c4", or e3b0c4 interchangeably. The isDocid utility then validates that the cleaned string contains at least six hexadecimal digits before processing.
Implementation Examples
The following patterns demonstrate how to generate and resolve docids using QMD's public API:
import { hashContent, getDocid } from "./src/store.ts";
async function makeDocid(text: string) {
const fullHash = await hashContent(text); // → "e3b0c44298fc1c149afbf4c8996fb924..."
const docid = getDocid(fullHash); // → "e3b0c4"
console.log(`#${docid}`); // prints "#e3b0c4"
}
For reverse lookup from a user-provided identifier:
import { findDocumentByDocid } from "./src/store.ts";
function getPathFromDocid(db, raw) {
const result = findDocumentByDocid(db, raw); // raw can be "#e3b0c4", "e3b0c4", etc.
if (result) {
console.log(`Virtual path: ${result.filepath}`);
} else {
console.log("No document matches that docid.");
}
}
Summary
- QMD's docid system generates identifiers using SHA-256 content hashing rather than filenames or timestamps.
- The
hashContentfunction insrc/store.tsproduces a full 64-character hexadecimal digest of the document body. - The
getDocidfunction extracts the first six characters to create a short, human-readable identifier presented as#xxxxxx. - The system stores only the full hash in the database, deriving the short docid on demand via
findDocumentByDocidusing SQLLIKEpattern matching. - Input normalization via
normalizeDocidallows flexible reference formats including#e3b0c4,"e3b0c4", or plaine3b0c4.
Frequently Asked Questions
What happens if two documents have the same 6-character docid prefix?
If multiple documents share the same six-character hash prefix, QMD's findDocumentByDocid function returns the first active match from the database query. The system uses LIMIT 1 in the SQL lookup, meaning collisions resolve to whichever document appears first in the index. For most practical collections, the 16.7 million possible six-character hex combinations provide sufficient entropy to avoid frequent collisions.
Why does QMD use SHA-256 instead of a shorter hash function?
QMD uses SHA-256 to ensure cryptographic uniqueness and future-proofing against collision attacks. While the public docid exposes only six characters for usability, the system stores the full 64-character digest internally. This design allows QMD to detect identical content with absolute certainty using the full hash while presenting users with a manageable short identifier. SHA-256 also provides consistent performance and widespread library support in Node.js environments.
How does QMD handle docid input formatting variations?
The normalizeDocid utility in src/store.ts strips optional formatting characters to accept multiple input styles. It removes surrounding double quotes and leading # characters, converting inputs like "e3b0c4" or #e3b0c4 into the canonical e3b0c4 format. The companion isDocid function then validates that the cleaned string contains at least six hexadecimal digits before the system processes the lookup, ensuring robust command-line and API interactions.
Does changing a single character in a document change the docid?
Yes, modifying even a single character produces an entirely new docid. Because QMD's hashContent function computes the SHA-256 digest of the complete document text, any alteration—whether adding a space, changing punctuation, or editing a word—results in a different 64-character hash and consequently a different 6-character docid prefix. This content-addressed approach guarantees that the docid always reflects the exact current state of the document.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →