Where to Find the Ingestion Pipeline Specification in LLM Wiki

The ingestion pipeline specification in LLM Wiki is distributed across TypeScript modules in src/lib/ and documented in README.md and plans/multimodal-images.md, with the main orchestration logic located in src/lib/ingest.ts.

The nashsu/llm_wiki repository transforms raw markdown and multimedia sources into fully-featured, vector-searchable wiki pages through a modular ingestion system. The complete ingestion pipeline specification spans implementation code, configuration files, and architectural design documents that collectively define how documents are parsed, enriched, chunked, and embedded.

Core Pipeline Implementation

The executable specification resides entirely within the src/lib/ directory, where each stage of the pipeline is implemented as a distinct module coordinated by a central driver.

Main Orchestration Driver

The pipeline entry point is src/lib/ingest.ts, which exports the autoIngest function. According to the source code at line 3194, autoIngest calls embedPage as the final step after completing parsing, captioning, and chunking operations. This module sequences the entire transformation from raw source document to searchable wiki page, handling error recovery and state management throughout the process.

Image Extraction and Captioning

Multimedia processing begins in src/lib/extract-source-images.ts, which discovers images within source documents and prepares them for downstream processing. The inline documentation identifies this as "Image extraction orchestration for the ingest pipeline."

Discovered images pass to src/lib/image-caption-pipeline.ts (explicitly referenced at line 132 in the codebase), which generates alt-text descriptions and caches results to avoid redundant API calls. This ensures all images in the final wiki pages include searchable, accessible descriptions.

Text Chunking and Vector Embedding

Text processing for Retrieval-Augmented Generation (RAG) occurs in two specialized modules:

  • src/lib/text-chunker.ts – Splits markdown content into semantically coherent chunks while preserving document structure.
  • src/lib/embedding.ts – Generates vector embeddings via the embedPage function, storing vectors in the search index for semantic retrieval.

These modules convert raw markdown into the vector representations that power the wiki's search functionality.

Configuration and Cleanup

Pipeline behavior is controlled by src/lib/source-watch-config.ts, which manages ingestion concurrency (defaulting to 1 parallel job) and file-watching parameters. After processing completes, src/lib/wiki-page-resolver.ts writes summary pages and updates front-matter metadata, while src/lib/wiki-cleanup.ts removes orphaned pages that no longer correspond to existing source files.

Documentation and Design Specifications

Beyond the implementation code, the ingestion pipeline specification appears in two critical design documents:

  • README.md – The "Auto-ingest" section provides the high-level architectural overview, explaining how clipped content triggers the two-stage pipeline.
  • plans/multimodal-images.md – Defines file-type routing logic, specifying which documents enter the full pipeline versus excluded standalone image files.

Programmatic Pipeline Invocation

Developers can interact with the pipeline specification directly through the exported functions:

// Manually invoking the ingest pipeline on a markdown file
import { autoIngest } from '@/lib/ingest';

const sourcePath = 'wiki/sources/example.md';
await autoIngest(sourcePath);
// Extracting images before captioning
import { extractImages } from '@/lib/extract-source-images';

const images = await extractImages('wiki/sources/example.md');
// Returns: Array<{ path: string, sha256: string }> ready for captioning
// Embedding a page after processing
import { embedPage } from '@/lib/embedding';

await embedPage('wiki/pages/example.md');

Summary

Frequently Asked Questions

What is the main entry point for the LLM Wiki ingestion pipeline?

The primary entry point is src/lib/ingest.ts, which exports the autoIngest function. This module coordinates all pipeline stages including parsing, image processing, chunking, and embedding, ultimately calling embedPage at line 3194 to finalize the vector representation.

How does LLM Wiki handle image processing during ingestion?

Images are discovered in src/lib/extract-source-images.ts, then processed by src/lib/image-caption-pipeline.ts to generate cached alt-text descriptions. This occurs early in the pipeline before text chunking, ensuring images are searchable within the vector store.

Where is the ingestion pipeline concurrency configured?

Concurrency settings are defined in src/lib/source-watch-config.ts, which defaults to 1 parallel ingestion job. This prevents API rate limit violations during embedding and caption generation operations.

How do I manually trigger the ingestion pipeline for a specific file?

Import autoIngest from @/lib/ingest and pass the source file path relative to the wiki directory. The function handles the complete lifecycle including extraction, captioning, chunking, embedding, and cleanup automatically.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →