What Is the Two-Step Chain-of-Thought Ingest Process in LLM Wiki?
The two-step chain-of-thought ingest process is an auto-ingest pipeline that separates document reasoning from content creation, using a structured analysis phase followed by a file generation phase to improve accuracy and traceability.
The nashsu/llm_wiki repository implements this innovative architecture to build knowledge bases with reduced hallucinations and higher consistency. By splitting the LLM's workflow into distinct cognitive stages, the system creates a reliable "scratchpad" that grounds final wiki pages in a coherent interpretation of source materials.
How the Two-Step Chain-of-Thought Ingest Process Works
The pipeline executes two sequential LLM invocations within src/lib/ingest.ts, each marked by explicit step comments that separate reasoning from output generation.
Step 1: Analysis Phase
The process begins with the LLM reading the raw source document and producing a structured analysis. At line 1024, the code marks this phase with the comment // Step 1: Analysis, during which the model identifies key entities, concepts, arguments, and contradictions with existing wiki content.
The analysis prompt is constructed from the project's configuration files—including purpose.md, schema.md, index.md, and overview.md—and transmitted to the model via the streamChat function at lines 1035-1042. This phase acts as the "scratchpad," allowing the model to establish a consistent interpretation before generating files.
Step 2: Generation Phase
Using the analysis as mandatory context, the pipeline enters the generation phase at line 1062, marked by // Step 2: Generation. During this stage, the LLM emits only FILE/REVIEW blocks, with each output starting with ---FILE: and omitting any preamble or explanatory text (lines 1068-1074).
This strict output format ensures that the generation phase remains focused solely on wiki file creation and review item generation, referencing the stable analysis context rather than re-interpreting the raw source.
Implementation in src/lib/ingest.ts
The core logic resides in the autoIngest function, which orchestrates both LLM calls while maintaining thread safety through a project-level mutex. The function accepts a project root, source file path, and LLM configuration, then executes the two-step chain-of-thought ingest sequentially.
import { autoIngest } from "./ingest"
import { useWikiStore } from "@/stores/wiki-store"
// Example: ingest a new PDF file
await autoIngest(
"/my/project", // project root
"/my/project/raw/sources/paper.pdf", // source file
useWikiStore.getState().llmConfig, // active LLM configuration
)
Internally, autoIngest first calls streamChat with the analysis prompt (lines 1035-1042), stores the structured reasoning, then invokes streamChat again with the generation prompt constrained to emit only ---FILE: blocks (lines 1068-1074).
Supporting Infrastructure and Safety Mechanisms
The two-step chain-of-thought ingest process relies on several auxiliary systems to ensure reliability:
src/lib/ingest-queue.ts– Serializes ingest tasks into a crash-safe queue, guaranteeing that multiple sources process sequentially without corrupting the wiki statesrc/lib/sweep-reviews.ts– Executes after Step 2 to resolve review items generated during the generation phase, creating a feedback loop for continuous improvement- Source hash caching – The pipeline computes and caches source file hashes, skipping unchanged files to avoid redundant LLM calls
- Optional image captioning – Runs before Step 1 when processing visual content, extracting text descriptions that feed into the analysis phase
Summary
- The two-step chain-of-thought ingest process deliberately separates reasoning (Step 1) from file generation (Step 2) to create a stable analytical scratchpad
- Step 1 analyzes documents using project context from
purpose.md,schema.md,index.md, andoverview.md, invoked viastreamChatat lines 1035-1042 ofsrc/lib/ingest.ts - Step 2 generates wiki files and review items constrained to
---FILE:blocks, implemented at lines 1068-1074 - The
autoIngestfunction orchestrates both steps within a project-level mutex to prevent race conditions - Supporting infrastructure includes queue management in
ingest-queue.tsand review resolution insweep-reviews.ts
Frequently Asked Questions
What triggers the two-step chain-of-thought ingest process?
The process triggers when autoIngest is called with a new or modified source file. The system first checks cached source hashes to skip unchanged files, optionally runs image captioning for visual content, then proceeds to Step 1 analysis if the content requires processing.
How does the analysis phase differ from the generation phase?
The analysis phase produces free-form structured reasoning about entities, concepts, and contradictions, while the generation phase is strictly constrained to emit only ---FILE: and review blocks. This separation ensures the model references a stable interpretation rather than hallucinating new analysis during file creation.
What files are involved in the ingest pipeline?
The primary implementation lives in src/lib/ingest.ts, with task serialization handled by src/lib/ingest-queue.ts and post-generation review management in src/lib/sweep-reviews.ts. Configuration files—including purpose.md, schema.md, index.md, and overview.md—provide context for the Step 1 prompt construction.
How does LLM Wiki prevent race conditions during ingest?
A project-level mutex wraps both LLM calls within autoIngest, ensuring serialized writes to the wiki knowledge base. This guarantees that the two-step chain-of-thought ingest process completes atomically for each source file, preventing corruption when multiple ingestion tasks run concurrently.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →