# How LLM Wiki Implements Incremental Caching for Ingestion

> Discover how LLM Wiki uses incremental caching with SHA-256 hashes to speed up ingestion, avoiding redundant LLM processing for unchanged files while updating image captions.

- Repository: [nash_su/llm_wiki](https://github.com/nashsu/llm_wiki)
- Tags: internals
- Published: 2026-09-13

---

**LLM Wiki implements incremental caching for ingestion by storing SHA-256 content hashes in a JSON cache file, skipping expensive LLM processing when source files remain unchanged while still refreshing image captions on cache hits.**

LLM Wiki, an open-source knowledge base generator by nashsu, optimizes repeated ingest operations through a deterministic caching layer. The system avoids redundant LLM calls by tracking source file fingerprints and generated outputs. Understanding this **incremental caching for ingestion** mechanism reveals how the project balances processing efficiency with content freshness.

## Core Cache Architecture

### Cache Storage Location and Format

The ingest cache persists as a JSON file located at `.<project-root>/.llm-wiki/ingest-cache.json`. This centralized store maps source file identities to their processing history, including content hashes and output file lists. The implementation in [`src/lib/ingest-cache.ts`](https://github.com/nashsu/llm_wiki/blob/main/src/lib/ingest-cache.ts) handles all read and write operations to this location.

### Content Hashing with SHA-256

Every source file receives a cryptographic fingerprint using SHA-256. The `sha256()` function computes a digest of the raw text content, storing this value as the `hash` field in the cache entry. This approach ensures that even minor content changes invalidate the cache, triggering a full re-ingest when necessary.

## Cache Hit Detection Logic

### The checkIngestCache Function

When the ingest pipeline initiates, it invokes `checkIngestCache(projectPath, sourceIdentity, sourceContent)` from [`src/lib/ingest.ts`](https://github.com/nashsu/llm_wiki/blob/main/src/lib/ingest.ts) around line 730. This utility loads the JSON cache and locates the entry keyed by the source filename.

```typescript
// src/lib/ingest.ts – around line 730
const cachedFiles = await checkIngestCache(projectRoot, sourceIdentity, sourceContent);
if (cachedFiles !== null) {
  // Cache hit – skip full ingest, just run image extraction
  await extractAndSaveSourceImages(projectRoot, sourcePath, sourceSlug);
  // …image captioning & injection…
  return; // done
}

// …full ingest pipeline runs when cache miss…

```

The function returns the list of previously generated files only if the stored hash matches the newly computed digest.

### File Integrity Verification

Cache hits require more than hash matching. The system validates that every file listed in the `filesWritten` array still exists on disk (see lines 80-96 in [`src/lib/ingest-cache.ts`](https://github.com/nashsu/llm_wiki/blob/main/src/lib/ingest-cache.ts)). Missing or deleted output files trigger cache invalidation, forcing a complete re-ingest to restore the wiki state. This safety check prevents stale cache entries from referencing non-existent content.

## Partial Pipeline Execution on Hits

Unlike naive caching systems, LLM Wiki does not skip all processing during a cache hit. While the heavy text-analysis and page-generation steps are bypassed, the pipeline still executes `extractAndSaveSourceImages()` starting at line 62 of [`src/lib/ingest.ts`](https://github.com/nashsu/llm_wiki/blob/main/src/lib/ingest.ts). This design ensures that newly-added image-captioning features or changed media files are incorporated without re-processing the entire source document.

## Cache Maintenance and Lifecycle

### Persisting New Cache Entries

After a successful full ingest, `saveIngestCache(projectPath, sourceFileName, sourceContent, filesWritten)` writes a fresh entry containing the current hash, timestamp, and complete list of generated wiki files. This operation in [`src/lib/ingest-cache.ts`](https://github.com/nashsu/llm_wiki/blob/main/src/lib/ingest-cache.ts) (lines 101-118) creates the checkpoint for future incremental runs.

```typescript
// src/lib/ingest.ts – after the pipeline writes wiki pages
const generatedFiles = [
  `wiki/${slug}.md`,
  `wiki/sources/${slug}.md`,
  // …any other files created during this ingest…
];
await saveIngestCache(projectRoot, sourceIdentity, sourceContent, generatedFiles);

```

### Handling Source File Renames

When users rename source documents, `moveIngestCacheEntry` (lines 34-51) preserves the processing history by copying the cache entry to a new key. The function accepts an optional path mapping to update file references within the `filesWritten` list.

```typescript
// src/lib/source-lifecycle.ts – when a source is moved
await moveIngestCacheEntry(
  projectRoot,
  oldSourceIdentity,
  newSourceIdentity,
  new Map([["old/path.md", "new/path.md"]])
);

```

### Cleaning Up Deleted Sources

The `removeFromIngestCache` function (lines 24-31) purges entries for deleted sources, preventing the cache file from accumulating orphaned references. Called from [`src/lib/source-lifecycle.ts`](https://github.com/nashsu/llm_wiki/blob/main/src/lib/source-lifecycle.ts), this cleanup ensures the cache remains an accurate index of current project state.

## Summary

- The ingest cache lives at `.<project-root>/.llm-wiki/ingest-cache.json` and stores SHA-256 content hashes alongside generated file lists.
- `checkIngestCache` validates both content hashes and file existence before declaring a cache hit, ensuring data integrity.
- Cache hits skip LLM processing but still execute image extraction pipelines, balancing speed with media freshness.
- Lifecycle functions `moveIngestCacheEntry` and `removeFromIngestCache` maintain cache consistency when sources are renamed or deleted.

## Frequently Asked Questions

### Where does LLM Wiki store its incremental ingest cache?

The cache persists as a JSON file at `.<project-root>/.llm-wiki/ingest-cache.json`, managed entirely by the functions in [`src/lib/ingest-cache.ts`](https://github.com/nashsu/llm_wiki/blob/main/src/lib/ingest-cache.ts). This location keeps cache data co-located with the project while remaining hidden from version control.

### How does LLM Wiki detect if a source file has changed?

The system computes a SHA-256 hash of the raw source content using the `sha256()` function. During ingest, `checkIngestCache` compares this fresh digest against the stored `hash` value. Any mismatch triggers a full re-ingest, ensuring the wiki reflects the latest source material.

### What processing still occurs when the cache hits?

Even on cache hits, LLM Wiki executes the image extraction and captioning sub-pipeline via `extractAndSaveSourceImages()`. This partial processing ensures new image-captioning features or modified media files are processed without the expense of re-analyzing the source text.

### How does the cache handle renamed or deleted source files?

The system provides `moveIngestCacheEntry` for renames, which transfers cache entries between keys while updating file paths, and `removeFromIngestCache` for deletions. These utilities, called from [`src/lib/source-lifecycle.ts`](https://github.com/nashsu/llm_wiki/blob/main/src/lib/source-lifecycle.ts), prevent stale entries and maintain cache accuracy across file system operations.