How Litho's Preprocessing Stage Analyzes Source Code: A Deep Dive into the deepwiki-rs Engine

Litho's preprocessing stage analyzes source code through a six-step pipeline orchestrated by PreProcessAgent, combining static file extraction with AI-driven code insights to generate structured artifacts for downstream documentation generation.

The Litho preprocessing stage serves as the critical first phase in the sopaco/deepwiki-rs documentation engine, transforming raw repositories into machine-readable structures. Operating as the entry point of the four-stage workflow (Preprocess → Research → Composition → Output), this stage extracts project metadata, identifies core source files, and generates AI-powered code insights that enable downstream agents to reason about architecture and relationships.

The Six-Step Preprocessing Pipeline

The preprocessing work is orchestrated by PreProcessAgent (defined in src/generator/preprocess/mod.rs). Its execute method implements a strict sequential pipeline that transforms raw code into structured intelligence.

Step 1: Extracting Original Documentation

The pipeline begins by calling original_document_extractor::extract to gather existing project documentation. This captures README files, changelogs, and other human-written materials that provide context for the AI analysis.

original_document_extractor::extract(&context).await?

Step 2: Building the Project Structure

StructureExtractor (located in src/generator/preprocess/extractors/structure_extractor.rs) performs a comprehensive filesystem scan. It counts files and directories, then constructs a ProjectStructure tree that maps the entire repository layout.

structure_extractor.extract_structure(&config.project_path).await?

Step 3: Identifying Core Source Files

The extractor applies heuristics to flag "important" modules based on file size, export count, and language-specific importance metrics. The identify_core_codes method filters for entry points like main.rs and public API surfaces.

structure_extractor.identify_core_codes(&project_structure).await?

Step 4: AI-Enhanced Code Analysis

The CodeAnalyze agent (implemented in src/generator/preprocess/agents/code_analyze.rs) sends core files to the LLM. It returns CodeInsight objects containing responsibilities, architectural roles, and complexity assessments.

code_analyze.execute(&context, &important_codes, &project_structure).await?

Step 5: Relationship Analysis

RelationshipsAnalyze (found in src/generator/preprocess/agents/relationships_analyze.rs) constructs a RelationshipAnalysis graph mapping calls, imports, and data flows between the previously generated CodeInsight objects.

relationships_analyze.execute(&context, &core_code_insights, &project_structure).await?

Step 6: Persisting Results to Memory

All artifacts—including OriginalDocument, ProjectStructure, CodeInsight collections, and RelationshipAnalysis—are stored in the Preprocess memory scope (MemoryScope::PREPROCESS) for downstream agent consumption.

context.memory().write(...)

Language-Agnostic Extraction Architecture

The static analysis component implements a language-processor plug-in system centered on the LanguageProcessor trait.

The LanguageProcessor Trait

Each supported language implements LanguageProcessor (see src/generator/preprocess/extractors/language_processors/php.rs for a concrete example). The trait exposes regex-based methods for extracting imports, exports, function signatures, and complexity metrics.

pub trait LanguageProcessor {
    fn file_extensions(&self) -> Vec<&'static str>;
    fn extract_imports(&self, content: &str) -> Vec<String>;
    // Additional extraction methods...
}

Asynchronous Processing of Core Files

LanguageProcessorManager (defined in src/generator/preprocess/extractors/language_processors/mod.rs) orchestrates all registered processors. These run asynchronously and are invoked only for files flagged as core (importance > 0.5), minimizing computational overhead.

Token-Efficient LLM Interaction

The AI phase implements two-round analysis with strict token management to handle large codebases efficiently.

Two-Round Analysis Strategy

  1. Raw syntactic extraction generates initial CodeInsight objects with fields like name, description, and importance_score.
  2. Architectural synthesis refines these insights into higher-level abstractions.

Prompt Compression and Limits

PromptCompressor (located in src/utils/prompt_compressor.rs) shrinks large insight collections to stay within token budgets while preserving architectural meaning. The pipeline enforces hard limits: top 150 insights per project and 20 dependencies per file.

Caching and Incremental Runs

StructureExtractor implements MD5-based caching of directory trees. If the hash of the project path matches a previous run, the static extraction phase is skipped entirely. This makes the preprocessing stage idempotent and enables fast incremental documentation builds, as detailed in docs/en/4.Deep-Exploration/Preprocessing Domain.md.

Downstream Consumption by Research Agents

Research agents retrieve preprocessing artifacts via MemoryScope::PREPROCESS. They consume MemoryData to read CodeInsight collections and RelationshipAnalysis, then perform higher-level architectural synthesis such as generating C4 diagrams and component maps.

Code Examples

Running the Full Pipeline via CLI


# Run Litho on a repository; the preprocessing stage executes automatically

deepwiki-rs generate --config litho-example.toml

Skipping Preprocessing with Cache

deepwiki-rs generate --config litho-example.toml --skip-preprocessing

The skip_preprocessing flag is defined in src/cli.rs.

Directly Invoking PreProcessAgent in Rust

use deepwiki_rs::generator::preprocess::PreProcessAgent;
use deepwiki_rs::generator::context::GeneratorContext;

#[tokio::main]
async fn main() -> anyhow::Result<()> {
    // Build a GeneratorContext (config, memory, llm client, etc.)
    let context = GeneratorContext::new("./my_project")?;

    // Create the agent and run it
    let preprocessor = PreProcessAgent::new();
    let result = preprocessor.execute(context).await?;

    // `result` now holds:
    // - result.original_document
    // - result.project_structure
    // - result.core_code_insights
    // - result.relationships
    println!("Preprocessing completed in {:.2}s", result.processing_time);
    Ok(())
}

Implementing a Custom Language Processor

// src/generator/preprocess/extractors/language_processors/my_lang.rs
use super::LanguageProcessor;

pub struct MyLangProcessor;

impl LanguageProcessor for MyLangProcessor {
    fn file_extensions(&self) -> Vec<&'static str> {
        vec!["my"]
    }

    fn extract_imports(&self, content: &str) -> Vec<String> {
        // Regex-based extraction here …
        vec![]
    }

    // Implement other required methods...
}

Then register it in LanguageProcessorManager (see src/generator/preprocess/extractors/language_processors/mod.rs).

Summary

  • Litho's preprocessing stage operates as the first phase of the four-stage deepwiki-rs workflow, transforming raw repositories into structured machine-readable artifacts.
  • The six-step pipeline (documentation extraction, structure building, core file identification, AI analysis, relationship mapping, and persistence) is orchestrated by PreProcessAgent in src/generator/preprocess/mod.rs.
  • Language-agnostic extraction relies on the LanguageProcessor trait and LanguageProcessorManager, enabling regex-based static analysis for any supported language.
  • Token efficiency is maintained through PromptCompressor and strict limits (150 insights per project, 20 dependencies per file) to manage LLM context windows.
  • Incremental builds are supported via MD5-based caching in StructureExtractor, making the stage idempotent and performant for repeated runs.

Frequently Asked Questions

What is the PreProcessAgent in Litho?

PreProcessAgent is the core orchestrator of the preprocessing stage in the deepwiki-rs engine. Defined in src/generator/preprocess/mod.rs, it implements a six-step sequential pipeline that extracts documentation, builds project structures, identifies core source files, generates AI-driven code insights, analyzes relationships, and persists all artifacts to the PREPROCESS memory scope for downstream research agents.

How does Litho identify which source files are important?

Litho uses the StructureExtractor (located in src/generator/preprocess/extractors/structure_extractor.rs) to apply heuristics that flag "core" files. The identify_core_codes method evaluates files based on size, export count, and language-specific importance metrics, filtering for entry points like main.rs and public API surfaces. Only files with an importance score greater than 0.5 are forwarded to the AI analysis phase.

Can Litho preprocess repositories incrementally?

Yes, Litho supports incremental preprocessing through MD5-based caching implemented in StructureExtractor. The system hashes the directory tree of the project path; if the hash matches a previous run, the static extraction phase is skipped entirely. This makes the preprocessing stage idempotent and significantly faster for repeated documentation builds, as detailed in the official documentation at docs/en/4.Deep-Exploration/Preprocessing Domain.md.

How does Litho handle different programming languages during preprocessing?

Litho implements a language-agnostic architecture through the LanguageProcessor trait and LanguageProcessorManager (found in src/generator/preprocess/extractors/language_processors/). Each supported language provides a concrete implementation (such as php.rs) that exposes regex-based methods for extracting imports, exports, function signatures, and complexity metrics. These processors run asynchronously and are invoked only for core files, allowing the system to support new languages by implementing the trait and registering the processor in the manager.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →