How LLM Wiki Uses Hashing for Efficient Ingest and Deduplication
LLM Wiki employs a multi-layered hashing strategy using FNV-1a and SHA-256 to eliminate redundant processing, detect file moves, and cache expensive LLM calls across the ingestion pipeline.
The open-source project nashsu/llm_wiki implements deterministic hashing at every stage of its data pipeline to minimize compute costs. By fingerprinting content before invoking embedding models or caption generators, the system ensures identical inputs trigger cache lookups rather than redundant API calls.
Content-Level Deduplication with hashTextHex
At the heart of the ingestion workflow lies hashTextHex, a 64-bit FNV-1a implementation defined in src/lib/ingest.ts. This function generates a fixed 16-character hexadecimal string representing the full text content of a source file.
When a source is processed, the pipeline computes hashTextHex(sourceContent) and compares the result against persistent storage in src/lib/ingest-cache.ts. A matching hash signals that the file content is unchanged, allowing the system to skip parsing, chunking, and embedding operations entirely.
// src/lib/ingest.ts
import { hashTextHex } from "./ingest";
const sourceContent = await readFile("notes.md");
const contentHash = hashTextHex(sourceContent); // 16-char hex
// Used as the key in ingest-cache.json to check for previous processing
Stable Source Identity with stableSlugHash
To handle file renames and URL changes without breaking downstream references, LLM Wiki uses stableSlugHash from src/lib/source-identity.ts. This 32-bit FNV-1a function creates a short, repeatable hash appended to human-readable slugs.
The format ${slug}--${hash} ensures that moving a file from docs/old-name.md to docs/new-name.md preserves the source identity as long as the content hash remains consistent. This mechanism prevents duplicate entries in the knowledge base when files are reorganized.
// src/lib/source-identity.ts
import { stableSlugHash } from "./source-identity";
const slug = "my-awesome-article";
const hash = stableSlugHash(slug); // → "k5z9"
const fullSlug = `${slug}--${hash}`; // → "my-awesome-article--k5z9"
Binary Asset Caching for Images
For image captioning workflows, cryptographic hashing ensures expensive vision-language model calls execute only once per unique image. In src/lib/image-caption-pipeline.ts, the system computes SHA-256 hashes of raw image bytes to generate cache keys.
The captionCacheKey function combines this hash with language parameters to create a deterministic lookup key. If the pipeline encounters the same image hash again, it retrieves the cached caption rather than invoking the LLM.
// src/lib/image-caption-pipeline.ts
import { sha256OfBase64 } from "./utils";
import { captionCacheKey } from "./image-caption-pipeline";
const imgBytes = await fetchImage(url);
const imgHash = await sha256OfBase64(imgBytes.base64);
const cacheKey = captionCacheKey(imgHash, "en");
// Skip LLM call if cache[cacheKey] exists
Detecting File Moves and Changes
The synchronization layer in src/lib/scheduled-import.ts and src/lib/project-file-sync.ts uses size-aware hashing to detect file mutations and moves efficiently. Files smaller than 32 KB receive fast 64-bit FNV-1a hashes, while larger files use SHA-256 to minimize collision risks.
During synchronization, the system stores hashBefore and hashAfter values in task metadata. When a deletion and creation event share the same hash, the pipeline interprets this as a file move rather than a delete-create cycle, avoiding redundant data copying.
// src/lib/project-file-sync.ts
if (task.kind === "created" && task.hashAfter) {
createdByHash.set(task.hashAfter, [...(createdByHash.get(task.hashAfter) ?? []), task]);
}
if (task.kind === "deleted" && task.hashBefore) {
deletedByHash.set(task.hashBefore, [...(deletedByHash.get(task.hashBefore) ?? []), task]);
}
// Matching hashes indicate a move, suppressing duplicate copy/delete operations
Persistent Cache Management
The src/lib/ingest-cache.ts module persists content hashes as SHA-256 digests in base64 encoding alongside generated file manifests. On subsequent ingestion runs, the system recomputes the source hash and compares it against the cached value.
A mismatch triggers a full re-ingest, while a match returns the previously generated files immediately. This content-addressable approach reduces incremental update times from minutes to milliseconds for unchanged documents.
Summary
- FNV-1a hashing provides fast, deterministic fingerprints for text content and source identifiers in
hashTextHexandstableSlugHash. - SHA-256 protects against collisions when caching binary assets like images or large files.
- Hash-based move detection in
project-file-sync.tseliminates redundant copy operations during file reorganization. - Content-addressable caching ensures LLM calls, embeddings, and captions execute only once per unique input.
Frequently Asked Questions
What hash algorithm does LLM Wiki use for text deduplication?
LLM Wiki uses 64-bit FNV-1a hashing for text deduplication via the hashTextHex function in src/lib/ingest.ts. This algorithm provides excellent distribution and speed for short to medium text strings while producing a consistent 16-character hexadecimal output suitable for cache keys.
How does the system detect when a file has been renamed rather than modified?
The system compares hashBefore and hashAfter values stored in synchronization task metadata within src/lib/project-file-sync.ts. When a deletion task and creation task share identical hashes, the pipeline treats this as a move operation, preserving the stable slug and avoiding redundant data processing.
Why does the image captioning pipeline use SHA-256 instead of FNV-1a?
The src/lib/image-caption-pipeline.ts module uses SHA-256 for image hashing because binary assets require collision-resistant cryptographic hashing. Given the computational expense of vision-language model inference, the negligible performance cost of SHA-256 is justified by the elimination of collision risks that could cause incorrect cache hits.
Where is the ingest cache physically stored?
The ingest cache resides in src/lib/ingest-cache.ts as a JSON file mapping source content hashes to their generated artifacts. This cache stores SHA-256 digests in base64 encoding alongside file manifests, enabling the system to skip unchanged sources during incremental ingestion runs.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →