# How QMD's Docid System Generates 6-Character Content Hashes

> Learn how QMD's docid system generates 6-character content hashes using SHA-256 for content-addressed storage without filenames or timestamps.

- Repository: [Tobias Lütke/qmd](https://github.com/tobi/qmd)
- Tags: internals
- Published: 2026-02-16

---

**QMD generates deterministic 6-character document identifiers by computing the SHA-256 hash of the entire file content and extracting the first six hexadecimal characters, enabling content-addressed storage without relying on filenames or timestamps.**

The `tobi/qmd` repository implements a content-addressed storage system where every document receives a stable, unique identifier derived solely from its text. Unlike traditional systems that rely on file paths or timestamps, QMD's docid system ensures that identical content always produces identical identifiers, while any modification—down to a single character—generates an entirely new hash. This approach centers on two core operations: computing a full SHA-256 digest and truncating it to a human-friendly 6-character prefix.

## How QMD's Docid System Works

The generation pipeline follows a strict two-step process implemented in [`src/store.ts`](https://github.com/tobi/qmd/blob/main/src/store.ts). First, the system computes a cryptographic hash of the entire document body. Second, it extracts a short, printable identifier from that hash for user interaction and database indexing.

### Content Hashing with SHA-256

The `hashContent` function in [`src/store.ts`](https://github.com/tobi/qmd/blob/main/src/store.ts) (lines 1919–1923) creates a standard SHA-256 digest of the raw document string:

```typescript
export async function hashContent(content: string): Promise<string> {
  const hash = createHash("sha256");
  hash.update(content);
  return hash.digest("hex");
}

```

This function returns the complete 64-character hexadecimal string representing the document's cryptographic fingerprint. The system stores this full hash in the database to guarantee uniqueness, while the shorter `docid` serves as a user-facing handle.

### Extracting the 6-Character Identifier

Once the full hash is computed, the `getDocid` function (lines 944–948) truncates it to the first six characters:

```typescript
export function getDocid(hash: string): string {
  return hash.slice(0, 6);
}

```

This produces values like `e3b0c4`, which QMD presents to users as `#e3b0c4`. The 6-character length provides a practical balance between readability and collision resistance for typical document collections.

## Database Storage and Retrieval

The docid system integrates tightly with QMD's SQLite storage layer. Rather than storing the short identifier separately, the system derives it on demand from the persisted full hash, ensuring consistency across all queries.

### Storing the Full Hash

When indexing a document, QMD inserts the 64-character SHA-256 hash into the `content` table and references it from the `documents` table. The short `docid` is never stored as a distinct column; instead, the system relies on the `LIKE` operator to match prefixes against the full hash column.

### Looking Up Documents by Docid

The `findDocumentByDocid` function (lines 1535–1548) implements reverse lookup using SQL pattern matching:

```typescript
export function findDocumentByDocid(db: Database, docid: string) {
  const shortHash = normalizeDocid(docid);
  return db.prepare(`
    SELECT 'qmd://' || d.collection || '/' || d.path AS filepath,
           d.hash
    FROM documents d
    WHERE d.hash LIKE ? AND d.active = 1
    LIMIT 1
  `).get(`${shortHash}%`);
}

```

This query appends a wildcard (`%`) to the 6-character prefix and searches for the first active document with a matching hash start. Collision handling is intentionally simple: if multiple documents share the same 6-character prefix, the query returns the first match.

### Normalizing User Input

To accommodate various user input formats, the `normalizeDocid` helper strips optional decorators. It removes surrounding quotes and leading `#` characters, allowing users to reference documents as `#e3b0c4`, `"e3b0c4"`, or `e3b0c4` interchangeably. The `isDocid` utility then validates that the cleaned string contains at least six hexadecimal digits before processing.

## Implementation Examples

The following patterns demonstrate how to generate and resolve docids using QMD's public API:

```typescript
import { hashContent, getDocid } from "./src/store.ts";

async function makeDocid(text: string) {
  const fullHash = await hashContent(text);   // → "e3b0c44298fc1c149afbf4c8996fb924..."
  const docid = getDocid(fullHash);           // → "e3b0c4"
  console.log(`#${docid}`);                   // prints "#e3b0c4"
}

```

For reverse lookup from a user-provided identifier:

```typescript
import { findDocumentByDocid } from "./src/store.ts";

function getPathFromDocid(db, raw) {
  const result = findDocumentByDocid(db, raw); // raw can be "#e3b0c4", "e3b0c4", etc.
  if (result) {
    console.log(`Virtual path: ${result.filepath}`);
  } else {
    console.log("No document matches that docid.");
  }
}

```

## Summary

- **QMD's docid system** generates identifiers using SHA-256 content hashing rather than filenames or timestamps.
- The `hashContent` function in [`src/store.ts`](https://github.com/tobi/qmd/blob/main/src/store.ts) produces a full 64-character hexadecimal digest of the document body.
- The `getDocid` function extracts the first six characters to create a short, human-readable identifier presented as `#xxxxxx`.
- The system stores only the full hash in the database, deriving the short docid on demand via `findDocumentByDocid` using SQL `LIKE` pattern matching.
- Input normalization via `normalizeDocid` allows flexible reference formats including `#e3b0c4`, `"e3b0c4"`, or plain `e3b0c4`.

## Frequently Asked Questions

### What happens if two documents have the same 6-character docid prefix?

If multiple documents share the same six-character hash prefix, QMD's `findDocumentByDocid` function returns the first active match from the database query. The system uses `LIMIT 1` in the SQL lookup, meaning collisions resolve to whichever document appears first in the index. For most practical collections, the 16.7 million possible six-character hex combinations provide sufficient entropy to avoid frequent collisions.

### Why does QMD use SHA-256 instead of a shorter hash function?

QMD uses SHA-256 to ensure cryptographic uniqueness and future-proofing against collision attacks. While the public `docid` exposes only six characters for usability, the system stores the full 64-character digest internally. This design allows QMD to detect identical content with absolute certainty using the full hash while presenting users with a manageable short identifier. SHA-256 also provides consistent performance and widespread library support in Node.js environments.

### How does QMD handle docid input formatting variations?

The `normalizeDocid` utility in [`src/store.ts`](https://github.com/tobi/qmd/blob/main/src/store.ts) strips optional formatting characters to accept multiple input styles. It removes surrounding double quotes and leading `#` characters, converting inputs like `"e3b0c4"` or `#e3b0c4` into the canonical `e3b0c4` format. The companion `isDocid` function then validates that the cleaned string contains at least six hexadecimal digits before the system processes the lookup, ensuring robust command-line and API interactions.

### Does changing a single character in a document change the docid?

Yes, modifying even a single character produces an entirely new docid. Because QMD's `hashContent` function computes the SHA-256 digest of the complete document text, any alteration—whether adding a space, changing punctuation, or editing a word—results in a different 64-character hash and consequently a different 6-character docid prefix. This content-addressed approach guarantees that the docid always reflects the exact current state of the document.