# How Litho's Preprocessing Stage Analyzes Source Code: A Deep Dive into the deepwiki-rs Engine

> Discover how Litho's preprocessing stage analyzes source code using the deepwiki-rs engine. Learn about its six-step AI-driven pipeline for structured documentation generation.

- Repository: [Sopaco/deepwiki-rs](https://github.com/sopaco/deepwiki-rs)
- Tags: deep-dive
- Published: 2026-02-16

---

**Litho's preprocessing stage analyzes source code through a six-step pipeline orchestrated by `PreProcessAgent`, combining static file extraction with AI-driven code insights to generate structured artifacts for downstream documentation generation.**

The **Litho preprocessing stage** serves as the critical first phase in the `sopaco/deepwiki-rs` documentation engine, transforming raw repositories into machine-readable structures. Operating as the entry point of the four-stage workflow (`Preprocess → Research → Composition → Output`), this stage extracts project metadata, identifies core source files, and generates AI-powered code insights that enable downstream agents to reason about architecture and relationships.

## The Six-Step Preprocessing Pipeline

The preprocessing work is orchestrated by `PreProcessAgent` (defined in [`src/generator/preprocess/mod.rs`](https://github.com/sopaco/deepwiki-rs/blob/main/src/generator/preprocess/mod.rs)). Its `execute` method implements a strict sequential pipeline that transforms raw code into structured intelligence.

### Step 1: Extracting Original Documentation

The pipeline begins by calling `original_document_extractor::extract` to gather existing project documentation. This captures README files, changelogs, and other human-written materials that provide context for the AI analysis.

```rust
original_document_extractor::extract(&context).await?

```

### Step 2: Building the Project Structure

`StructureExtractor` (located in [`src/generator/preprocess/extractors/structure_extractor.rs`](https://github.com/sopaco/deepwiki-rs/blob/main/src/generator/preprocess/extractors/structure_extractor.rs)) performs a comprehensive filesystem scan. It counts files and directories, then constructs a `ProjectStructure` tree that maps the entire repository layout.

```rust
structure_extractor.extract_structure(&config.project_path).await?

```

### Step 3: Identifying Core Source Files

The extractor applies heuristics to flag "important" modules based on file size, export count, and language-specific importance metrics. The `identify_core_codes` method filters for entry points like [`main.rs`](https://github.com/sopaco/deepwiki-rs/blob/main/main.rs) and public API surfaces.

```rust
structure_extractor.identify_core_codes(&project_structure).await?

```

### Step 4: AI-Enhanced Code Analysis

The `CodeAnalyze` agent (implemented in [`src/generator/preprocess/agents/code_analyze.rs`](https://github.com/sopaco/deepwiki-rs/blob/main/src/generator/preprocess/agents/code_analyze.rs)) sends core files to the LLM. It returns `CodeInsight` objects containing responsibilities, architectural roles, and complexity assessments.

```rust
code_analyze.execute(&context, &important_codes, &project_structure).await?

```

### Step 5: Relationship Analysis

`RelationshipsAnalyze` (found in [`src/generator/preprocess/agents/relationships_analyze.rs`](https://github.com/sopaco/deepwiki-rs/blob/main/src/generator/preprocess/agents/relationships_analyze.rs)) constructs a `RelationshipAnalysis` graph mapping calls, imports, and data flows between the previously generated `CodeInsight` objects.

```rust
relationships_analyze.execute(&context, &core_code_insights, &project_structure).await?

```

### Step 6: Persisting Results to Memory

All artifacts—including `OriginalDocument`, `ProjectStructure`, `CodeInsight` collections, and `RelationshipAnalysis`—are stored in the **Preprocess memory scope** (`MemoryScope::PREPROCESS`) for downstream agent consumption.

```rust
context.memory().write(...)

```

## Language-Agnostic Extraction Architecture

The static analysis component implements a **language-processor plug-in system** centered on the `LanguageProcessor` trait.

### The LanguageProcessor Trait

Each supported language implements `LanguageProcessor` (see [`src/generator/preprocess/extractors/language_processors/php.rs`](https://github.com/sopaco/deepwiki-rs/blob/main/src/generator/preprocess/extractors/language_processors/php.rs) for a concrete example). The trait exposes regex-based methods for extracting imports, exports, function signatures, and complexity metrics.

```rust
pub trait LanguageProcessor {
    fn file_extensions(&self) -> Vec<&'static str>;
    fn extract_imports(&self, content: &str) -> Vec<String>;
    // Additional extraction methods...
}

```

### Asynchronous Processing of Core Files

`LanguageProcessorManager` (defined in [`src/generator/preprocess/extractors/language_processors/mod.rs`](https://github.com/sopaco/deepwiki-rs/blob/main/src/generator/preprocess/extractors/language_processors/mod.rs)) orchestrates all registered processors. These run **asynchronously** and are invoked only for files flagged as *core* (importance > 0.5), minimizing computational overhead.

## Token-Efficient LLM Interaction

The AI phase implements **two-round analysis** with strict token management to handle large codebases efficiently.

### Two-Round Analysis Strategy

1. **Raw syntactic extraction** generates initial `CodeInsight` objects with fields like `name`, `description`, and `importance_score`.
2. **Architectural synthesis** refines these insights into higher-level abstractions.

### Prompt Compression and Limits

`PromptCompressor` (located in [`src/utils/prompt_compressor.rs`](https://github.com/sopaco/deepwiki-rs/blob/main/src/utils/prompt_compressor.rs)) shrinks large insight collections to stay within token budgets while preserving architectural meaning. The pipeline enforces hard limits: **top 150 insights per project** and **20 dependencies per file**.

## Caching and Incremental Runs

`StructureExtractor` implements MD5-based caching of directory trees. If the hash of the project path matches a previous run, the static extraction phase is skipped entirely. This makes the preprocessing stage **idempotent** and enables fast incremental documentation builds, as detailed in `docs/en/4.Deep-Exploration/Preprocessing Domain.md`.

## Downstream Consumption by Research Agents

Research agents retrieve preprocessing artifacts via `MemoryScope::PREPROCESS`. They consume `MemoryData` to read `CodeInsight` collections and `RelationshipAnalysis`, then perform higher-level architectural synthesis such as generating C4 diagrams and component maps.

## Code Examples

### Running the Full Pipeline via CLI

```bash

# Run Litho on a repository; the preprocessing stage executes automatically

deepwiki-rs generate --config litho-example.toml

```

### Skipping Preprocessing with Cache

```bash
deepwiki-rs generate --config litho-example.toml --skip-preprocessing

```

The `skip_preprocessing` flag is defined in [`src/cli.rs`](https://github.com/sopaco/deepwiki-rs/blob/main/src/cli.rs).

### Directly Invoking PreProcessAgent in Rust

```rust
use deepwiki_rs::generator::preprocess::PreProcessAgent;
use deepwiki_rs::generator::context::GeneratorContext;

#[tokio::main]
async fn main() -> anyhow::Result<()> {
    // Build a GeneratorContext (config, memory, llm client, etc.)
    let context = GeneratorContext::new("./my_project")?;

    // Create the agent and run it
    let preprocessor = PreProcessAgent::new();
    let result = preprocessor.execute(context).await?;

    // `result` now holds:
    // - result.original_document
    // - result.project_structure
    // - result.core_code_insights
    // - result.relationships
    println!("Preprocessing completed in {:.2}s", result.processing_time);
    Ok(())
}

```

### Implementing a Custom Language Processor

```rust
// src/generator/preprocess/extractors/language_processors/my_lang.rs
use super::LanguageProcessor;

pub struct MyLangProcessor;

impl LanguageProcessor for MyLangProcessor {
    fn file_extensions(&self) -> Vec<&'static str> {
        vec!["my"]
    }

    fn extract_imports(&self, content: &str) -> Vec<String> {
        // Regex-based extraction here …
        vec![]
    }

    // Implement other required methods...
}

```

Then register it in `LanguageProcessorManager` (see [`src/generator/preprocess/extractors/language_processors/mod.rs`](https://github.com/sopaco/deepwiki-rs/blob/main/src/generator/preprocess/extractors/language_processors/mod.rs)).

## Summary

- **Litho's preprocessing stage** operates as the first phase of the four-stage `deepwiki-rs` workflow, transforming raw repositories into structured machine-readable artifacts.
- The **six-step pipeline** (documentation extraction, structure building, core file identification, AI analysis, relationship mapping, and persistence) is orchestrated by `PreProcessAgent` in [`src/generator/preprocess/mod.rs`](https://github.com/sopaco/deepwiki-rs/blob/main/src/generator/preprocess/mod.rs).
- **Language-agnostic extraction** relies on the `LanguageProcessor` trait and `LanguageProcessorManager`, enabling regex-based static analysis for any supported language.
- **Token efficiency** is maintained through `PromptCompressor` and strict limits (150 insights per project, 20 dependencies per file) to manage LLM context windows.
- **Incremental builds** are supported via MD5-based caching in `StructureExtractor`, making the stage idempotent and performant for repeated runs.

## Frequently Asked Questions

### What is the PreProcessAgent in Litho?

`PreProcessAgent` is the core orchestrator of the preprocessing stage in the `deepwiki-rs` engine. Defined in [`src/generator/preprocess/mod.rs`](https://github.com/sopaco/deepwiki-rs/blob/main/src/generator/preprocess/mod.rs), it implements a six-step sequential pipeline that extracts documentation, builds project structures, identifies core source files, generates AI-driven code insights, analyzes relationships, and persists all artifacts to the `PREPROCESS` memory scope for downstream research agents.

### How does Litho identify which source files are important?

Litho uses the `StructureExtractor` (located in [`src/generator/preprocess/extractors/structure_extractor.rs`](https://github.com/sopaco/deepwiki-rs/blob/main/src/generator/preprocess/extractors/structure_extractor.rs)) to apply heuristics that flag "core" files. The `identify_core_codes` method evaluates files based on size, export count, and language-specific importance metrics, filtering for entry points like [`main.rs`](https://github.com/sopaco/deepwiki-rs/blob/main/main.rs) and public API surfaces. Only files with an importance score greater than 0.5 are forwarded to the AI analysis phase.

### Can Litho preprocess repositories incrementally?

Yes, Litho supports incremental preprocessing through MD5-based caching implemented in `StructureExtractor`. The system hashes the directory tree of the project path; if the hash matches a previous run, the static extraction phase is skipped entirely. This makes the preprocessing stage idempotent and significantly faster for repeated documentation builds, as detailed in the official documentation at `docs/en/4.Deep-Exploration/Preprocessing Domain.md`.

### How does Litho handle different programming languages during preprocessing?

Litho implements a language-agnostic architecture through the `LanguageProcessor` trait and `LanguageProcessorManager` (found in `src/generator/preprocess/extractors/language_processors/`). Each supported language provides a concrete implementation (such as [`php.rs`](https://github.com/sopaco/deepwiki-rs/blob/main/php.rs)) that exposes regex-based methods for extracting imports, exports, function signatures, and complexity metrics. These processors run asynchronously and are invoked only for core files, allowing the system to support new languages by implementing the trait and registering the processor in the manager.