How LLM Wiki Implements Incremental Caching for Ingestion

LLM Wiki implements incremental caching for ingestion by storing SHA-256 content hashes in a JSON cache file, skipping expensive LLM processing when source files remain unchanged while still refreshing image captions on cache hits.

LLM Wiki, an open-source knowledge base generator by nashsu, optimizes repeated ingest operations through a deterministic caching layer. The system avoids redundant LLM calls by tracking source file fingerprints and generated outputs. Understanding this incremental caching for ingestion mechanism reveals how the project balances processing efficiency with content freshness.

Core Cache Architecture

Cache Storage Location and Format

The ingest cache persists as a JSON file located at .<project-root>/.llm-wiki/ingest-cache.json. This centralized store maps source file identities to their processing history, including content hashes and output file lists. The implementation in src/lib/ingest-cache.ts handles all read and write operations to this location.

Content Hashing with SHA-256

Every source file receives a cryptographic fingerprint using SHA-256. The sha256() function computes a digest of the raw text content, storing this value as the hash field in the cache entry. This approach ensures that even minor content changes invalidate the cache, triggering a full re-ingest when necessary.

Cache Hit Detection Logic

The checkIngestCache Function

When the ingest pipeline initiates, it invokes checkIngestCache(projectPath, sourceIdentity, sourceContent) from src/lib/ingest.ts around line 730. This utility loads the JSON cache and locates the entry keyed by the source filename.

// src/lib/ingest.ts – around line 730
const cachedFiles = await checkIngestCache(projectRoot, sourceIdentity, sourceContent);
if (cachedFiles !== null) {
  // Cache hit – skip full ingest, just run image extraction
  await extractAndSaveSourceImages(projectRoot, sourcePath, sourceSlug);
  // …image captioning & injection…
  return; // done
}

// …full ingest pipeline runs when cache miss…

The function returns the list of previously generated files only if the stored hash matches the newly computed digest.

File Integrity Verification

Cache hits require more than hash matching. The system validates that every file listed in the filesWritten array still exists on disk (see lines 80-96 in src/lib/ingest-cache.ts). Missing or deleted output files trigger cache invalidation, forcing a complete re-ingest to restore the wiki state. This safety check prevents stale cache entries from referencing non-existent content.

Partial Pipeline Execution on Hits

Unlike naive caching systems, LLM Wiki does not skip all processing during a cache hit. While the heavy text-analysis and page-generation steps are bypassed, the pipeline still executes extractAndSaveSourceImages() starting at line 62 of src/lib/ingest.ts. This design ensures that newly-added image-captioning features or changed media files are incorporated without re-processing the entire source document.

Cache Maintenance and Lifecycle

Persisting New Cache Entries

After a successful full ingest, saveIngestCache(projectPath, sourceFileName, sourceContent, filesWritten) writes a fresh entry containing the current hash, timestamp, and complete list of generated wiki files. This operation in src/lib/ingest-cache.ts (lines 101-118) creates the checkpoint for future incremental runs.

// src/lib/ingest.ts – after the pipeline writes wiki pages
const generatedFiles = [
  `wiki/${slug}.md`,
  `wiki/sources/${slug}.md`,
  // …any other files created during this ingest…
];
await saveIngestCache(projectRoot, sourceIdentity, sourceContent, generatedFiles);

Handling Source File Renames

When users rename source documents, moveIngestCacheEntry (lines 34-51) preserves the processing history by copying the cache entry to a new key. The function accepts an optional path mapping to update file references within the filesWritten list.

// src/lib/source-lifecycle.ts – when a source is moved
await moveIngestCacheEntry(
  projectRoot,
  oldSourceIdentity,
  newSourceIdentity,
  new Map([["old/path.md", "new/path.md"]])
);

Cleaning Up Deleted Sources

The removeFromIngestCache function (lines 24-31) purges entries for deleted sources, preventing the cache file from accumulating orphaned references. Called from src/lib/source-lifecycle.ts, this cleanup ensures the cache remains an accurate index of current project state.

Summary

  • The ingest cache lives at .<project-root>/.llm-wiki/ingest-cache.json and stores SHA-256 content hashes alongside generated file lists.
  • checkIngestCache validates both content hashes and file existence before declaring a cache hit, ensuring data integrity.
  • Cache hits skip LLM processing but still execute image extraction pipelines, balancing speed with media freshness.
  • Lifecycle functions moveIngestCacheEntry and removeFromIngestCache maintain cache consistency when sources are renamed or deleted.

Frequently Asked Questions

Where does LLM Wiki store its incremental ingest cache?

The cache persists as a JSON file at .<project-root>/.llm-wiki/ingest-cache.json, managed entirely by the functions in src/lib/ingest-cache.ts. This location keeps cache data co-located with the project while remaining hidden from version control.

How does LLM Wiki detect if a source file has changed?

The system computes a SHA-256 hash of the raw source content using the sha256() function. During ingest, checkIngestCache compares this fresh digest against the stored hash value. Any mismatch triggers a full re-ingest, ensuring the wiki reflects the latest source material.

What processing still occurs when the cache hits?

Even on cache hits, LLM Wiki executes the image extraction and captioning sub-pipeline via extractAndSaveSourceImages(). This partial processing ensures new image-captioning features or modified media files are processed without the expense of re-analyzing the source text.

How does the cache handle renamed or deleted source files?

The system provides moveIngestCacheEntry for renames, which transfers cache entries between keys while updating file paths, and removeFromIngestCache for deletions. These utilities, called from src/lib/source-lifecycle.ts, prevent stale entries and maintain cache accuracy across file system operations.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →