# How LLM Wiki Uses Hashing for Efficient Ingest and Deduplication

> Discover how LLM Wiki uses FNV-1a and SHA-256 hashing for efficient ingest and deduplication, eliminating redundant processing and caching LLM calls.

- Repository: [nash_su/llm_wiki](https://github.com/nashsu/llm_wiki)
- Tags: internals
- Published: 2026-09-12

---

**LLM Wiki employs a multi-layered hashing strategy using FNV-1a and SHA-256 to eliminate redundant processing, detect file moves, and cache expensive LLM calls across the ingestion pipeline.**

The open-source project `nashsu/llm_wiki` implements deterministic hashing at every stage of its data pipeline to minimize compute costs. By fingerprinting content before invoking embedding models or caption generators, the system ensures identical inputs trigger cache lookups rather than redundant API calls.

## Content-Level Deduplication with `hashTextHex`

At the heart of the ingestion workflow lies `hashTextHex`, a 64-bit FNV-1a implementation defined in [`src/lib/ingest.ts`](https://github.com/nashsu/llm_wiki/blob/main/src/lib/ingest.ts). This function generates a fixed 16-character hexadecimal string representing the full text content of a source file.

When a source is processed, the pipeline computes `hashTextHex(sourceContent)` and compares the result against persistent storage in [`src/lib/ingest-cache.ts`](https://github.com/nashsu/llm_wiki/blob/main/src/lib/ingest-cache.ts). A matching hash signals that the file content is unchanged, allowing the system to skip parsing, chunking, and embedding operations entirely.

```typescript
// src/lib/ingest.ts
import { hashTextHex } from "./ingest";

const sourceContent = await readFile("notes.md");
const contentHash = hashTextHex(sourceContent);   // 16-char hex
// Used as the key in ingest-cache.json to check for previous processing

```

## Stable Source Identity with `stableSlugHash`

To handle file renames and URL changes without breaking downstream references, LLM Wiki uses `stableSlugHash` from [`src/lib/source-identity.ts`](https://github.com/nashsu/llm_wiki/blob/main/src/lib/source-identity.ts). This 32-bit FNV-1a function creates a short, repeatable hash appended to human-readable slugs.

The format `${slug}--${hash}` ensures that moving a file from [`docs/old-name.md`](https://github.com/nashsu/llm_wiki/blob/main/docs/old-name.md) to [`docs/new-name.md`](https://github.com/nashsu/llm_wiki/blob/main/docs/new-name.md) preserves the source identity as long as the content hash remains consistent. This mechanism prevents duplicate entries in the knowledge base when files are reorganized.

```typescript
// src/lib/source-identity.ts
import { stableSlugHash } from "./source-identity";

const slug = "my-awesome-article";
const hash = stableSlugHash(slug);          // → "k5z9"
const fullSlug = `${slug}--${hash}`;        // → "my-awesome-article--k5z9"

```

## Binary Asset Caching for Images

For image captioning workflows, cryptographic hashing ensures expensive vision-language model calls execute only once per unique image. In [`src/lib/image-caption-pipeline.ts`](https://github.com/nashsu/llm_wiki/blob/main/src/lib/image-caption-pipeline.ts), the system computes SHA-256 hashes of raw image bytes to generate cache keys.

The `captionCacheKey` function combines this hash with language parameters to create a deterministic lookup key. If the pipeline encounters the same image hash again, it retrieves the cached caption rather than invoking the LLM.

```typescript
// src/lib/image-caption-pipeline.ts
import { sha256OfBase64 } from "./utils";
import { captionCacheKey } from "./image-caption-pipeline";

const imgBytes = await fetchImage(url);
const imgHash = await sha256OfBase64(imgBytes.base64);
const cacheKey = captionCacheKey(imgHash, "en");

// Skip LLM call if cache[cacheKey] exists

```

## Detecting File Moves and Changes

The synchronization layer in [`src/lib/scheduled-import.ts`](https://github.com/nashsu/llm_wiki/blob/main/src/lib/scheduled-import.ts) and [`src/lib/project-file-sync.ts`](https://github.com/nashsu/llm_wiki/blob/main/src/lib/project-file-sync.ts) uses size-aware hashing to detect file mutations and moves efficiently. Files smaller than 32 KB receive fast 64-bit FNV-1a hashes, while larger files use SHA-256 to minimize collision risks.

During synchronization, the system stores `hashBefore` and `hashAfter` values in task metadata. When a deletion and creation event share the same hash, the pipeline interprets this as a file move rather than a delete-create cycle, avoiding redundant data copying.

```typescript
// src/lib/project-file-sync.ts
if (task.kind === "created" && task.hashAfter) {
  createdByHash.set(task.hashAfter, [...(createdByHash.get(task.hashAfter) ?? []), task]);
}
if (task.kind === "deleted" && task.hashBefore) {
  deletedByHash.set(task.hashBefore, [...(deletedByHash.get(task.hashBefore) ?? []), task]);
}

// Matching hashes indicate a move, suppressing duplicate copy/delete operations

```

## Persistent Cache Management

The [`src/lib/ingest-cache.ts`](https://github.com/nashsu/llm_wiki/blob/main/src/lib/ingest-cache.ts) module persists content hashes as SHA-256 digests in base64 encoding alongside generated file manifests. On subsequent ingestion runs, the system recomputes the source hash and compares it against the cached value.

A mismatch triggers a full re-ingest, while a match returns the previously generated files immediately. This content-addressable approach reduces incremental update times from minutes to milliseconds for unchanged documents.

## Summary

- **FNV-1a hashing** provides fast, deterministic fingerprints for text content and source identifiers in `hashTextHex` and `stableSlugHash`.
- **SHA-256** protects against collisions when caching binary assets like images or large files.
- **Hash-based move detection** in [`project-file-sync.ts`](https://github.com/nashsu/llm_wiki/blob/main/project-file-sync.ts) eliminates redundant copy operations during file reorganization.
- **Content-addressable caching** ensures LLM calls, embeddings, and captions execute only once per unique input.

## Frequently Asked Questions

### What hash algorithm does LLM Wiki use for text deduplication?

LLM Wiki uses 64-bit FNV-1a hashing for text deduplication via the `hashTextHex` function in [`src/lib/ingest.ts`](https://github.com/nashsu/llm_wiki/blob/main/src/lib/ingest.ts). This algorithm provides excellent distribution and speed for short to medium text strings while producing a consistent 16-character hexadecimal output suitable for cache keys.

### How does the system detect when a file has been renamed rather than modified?

The system compares `hashBefore` and `hashAfter` values stored in synchronization task metadata within [`src/lib/project-file-sync.ts`](https://github.com/nashsu/llm_wiki/blob/main/src/lib/project-file-sync.ts). When a deletion task and creation task share identical hashes, the pipeline treats this as a move operation, preserving the stable slug and avoiding redundant data processing.

### Why does the image captioning pipeline use SHA-256 instead of FNV-1a?

The [`src/lib/image-caption-pipeline.ts`](https://github.com/nashsu/llm_wiki/blob/main/src/lib/image-caption-pipeline.ts) module uses SHA-256 for image hashing because binary assets require collision-resistant cryptographic hashing. Given the computational expense of vision-language model inference, the negligible performance cost of SHA-256 is justified by the elimination of collision risks that could cause incorrect cache hits.

### Where is the ingest cache physically stored?

The ingest cache resides in [`src/lib/ingest-cache.ts`](https://github.com/nashsu/llm_wiki/blob/main/src/lib/ingest-cache.ts) as a JSON file mapping source content hashes to their generated artifacts. This cache stores SHA-256 digests in base64 encoding alongside file manifests, enabling the system to skip unchanged sources during incremental ingestion runs.