How Litho's Intelligent Research and Analysis Stage Works in deepwiki-rs

Litho's intelligent research and analysis stage uses a multi-agent pipeline orchestrated by ResearchOrchestrator to automatically extract architectural knowledge from codebases through five deterministic layers: macro context, domain architecture, workflows, key modules, and boundary analysis.

The deepwiki-rs engine (codenamed Litho) generates comprehensive documentation by first executing an intelligent research phase that explores a project's structure without manual annotation. This stage is implemented in the sopaco/deepwiki-rs repository and operates through a series of specialized agents that analyze code at increasing levels of granularity.

The Multi-Agent Research Pipeline

At the heart of Litho's intelligent research and analysis stage lies a deterministic pipeline where each agent builds upon the findings of its predecessors. The architecture follows a layered approach inspired by the C4 model, progressing from macro system context down to micro-level implementation details.

ResearchOrchestrator as the Central Controller

The ResearchOrchestrator struct in src/generator/research/orchestrator.rs serves as the central dispatcher that coordinates all research activities. It implements the execute_research_pipeline method, which sequences agent execution and manages data dependencies between layers.

// src/generator/workflow.rs – launch of the research stage
let research_orchestrator = ResearchOrchestrator::default();
research_orchestrator
    .execute_research_pipeline(&context)
    .await?;

The Five-Layer Analysis Strategy

The pipeline executes agents across five distinct analytical layers:

  1. C1 Macro – Captures overall purpose, business value, and system boundaries
  2. C2 Meso – Identifies domain divisions, architecture, and core workflows
  3. C3-C4 Micro – Drills into critical modules and implementation details
  4. Boundary Analysis – Enumerates external entry points and integration surfaces
  5. Database Overview – Models SQL schemas and data flows when applicable

Layer-by-Agent Breakdown

Each layer in Litho's intelligent research and analysis stage is handled by specialized agents that implement the StepForwardAgent trait, defining their own prompt templates, data requirements, and output schemas.

C1 Macro Layer – System Context

The SystemContextResearcher initiates the pipeline by analyzing project objectives, target users, and external system dependencies. Located in src/generator/research/agents/system_context_researcher.rs, this agent defines the system's contextual boundaries and establishes the foundation for subsequent analysis.

C2 Meso Layer – Domain and Architecture

Three agents collaborate to map the intermediate architectural layer:

DomainModulesDetector (src/generator/research/agents/domain_modules_detector.rs) identifies top-level business domains, sub-modules, and inter-domain relationships.

ArchitectureResearcher (src/generator/research/agents/architecture_researcher.rs) summarizes the system's structural layout, including components, layers, and deployment topology, producing Mermaid diagrams for visual representation.

WorkflowResearcher (src/generator/research/agents/workflow_researcher.rs) extracts main functional workflows, mapping process steps and business flows from code analysis and previous research reports.

C3-C4 Micro Layer – Key Module Insights

The KeyModulesInsight agent (src/generator/research/agents/key_modules_insight.rs) performs parallel micro-analysis of each domain discovered in previous layers. For every critical module, it generates a detailed KeyModuleReport containing:

  • Module description and responsibilities
  • Interface definitions
  • Implementation code snippets
  • Flow charts and sequence diagrams

This agent utilizes do_parallel_with_limit to respect the max_parallels configuration setting, ensuring efficient resource utilization during deep analysis.

Boundary and Database Analysis

BoundaryAnalyzer (src/generator/research/agents/boundary_analyzer.rs) scans code insights for entry-point purposes including Entry, Api, Router, and Config types. It returns a structured JSON report enumerating all external integration surfaces and suggesting potential third-party integrations.

DatabaseOverviewAnalyzer (src/generator/research/agents/database_overview_analyzer.rs) conditionally executes when SQL artifacts are present in the codebase. This agent filters SQL-related insights to extract tables, views, stored procedures, and relationships, producing a concise database schema overview.

Execution Flow and Data Management

The intelligent research and analysis stage implements sophisticated prompt engineering and memory management to ensure coherent, contextually aware analysis across all agents.

Prompt Engineering and LLM Integration

Each agent implements the StepForwardAgent trait, defining a prompt_template() method that returns a PromptTemplate struct containing:

  • System prompt – Defines the agent's persona and expertise
  • Opening/Closing instructions – Frames the specific analysis task
  • LLM call mode – Either Extract (for JSON-typed results) or PromptWithTools (for richer artifacts like Mermaid diagrams)
  • Formatter configuration – Controls output formatting

The ResearchOrchestrator resolves localized agent names via AgentType::display_name for clear console output, then invokes agent.execute(context).await. Internally, this builds the complete prompt and sends it to the configured LLM via LLMClient.

Parallel Processing and Memory Scopes

Research results are stored in the research memory scope (MemoryScope::STUDIES_RESEARCH), making them available to downstream agents through DataSource::ResearchResult. The orchestrator guarantees execution order such that prerequisite agents complete before dependent agents begin.

The KeyModulesInsight agent leverages parallel execution via do_parallel_with_limit, respecting the max_parallels setting from the global configuration to prevent overwhelming the LLM API while maximizing throughput.

Implementing Custom Research Agents

Extending Litho's intelligent research and analysis stage requires implementing the StepForwardAgent trait and registering the new agent in the orchestrator:

// src/generator/research/agents/custom_security_researcher.rs
#[derive(Default)]
pub struct SecurityResearcher;

impl StepForwardAgent for SecurityResearcher {
    type Output = SecurityReport;

    fn agent_type(&self) -> String { AgentType::SecurityResearcher.to_string() }
    fn memory_scope_key(&self) -> String { MemoryScope::STUDIES_RESEARCH.to_string() }

    fn data_config(&self) -> AgentDataConfig {
        AgentDataConfig {
            required_sources: vec![DataSource::CODE_INSIGHTS],
            optional_sources: vec![],
        }
    }

    fn prompt_template(&self) -> PromptTemplate {
        PromptTemplate {
            system_prompt: "You are a security analyst...".into(),
            opening_instruction: "Analyze the following code for security risks:".into(),
            closing_instruction: "".into(),
            llm_call_mode: LLMCallMode::Extract,
            formatter_config: FormatterConfig::default(),
        }
    }
}

To activate the custom agent, add the execution call in ResearchOrchestrator::execute_research_pipeline within src/generator/research/orchestrator.rs:

self.execute_agent(&SecurityResearcher::default(), context).await?;

Summary

  • Litho's intelligent research and analysis stage operates through a deterministic multi-agent pipeline orchestrated by ResearchOrchestrator in src/generator/research/orchestrator.rs.
  • The pipeline progresses through five analytical layers: macro system context, meso domain architecture, micro module details, boundary analysis, and optional database overview.
  • Each specialized agent implements the StepForwardAgent trait, defining custom prompt templates, data dependencies, and output schemas stored in the research memory scope.
  • Parallel execution is utilized for intensive micro-analysis via KeyModulesInsight, while the orchestrator ensures prerequisite data is available before dependent agents execute.
  • Developers can extend the pipeline by implementing custom agents and registering them in the orchestrator's execution sequence.

Frequently Asked Questions

What triggers the intelligent research and analysis stage in Litho?

The research stage is triggered when the main documentation workflow begins via launch() in src/generator/workflow.rs. This function instantiates ResearchOrchestrator and calls execute_research_pipeline(), which sequentially runs all research agents. The stage executes automatically as part of the standard documentation generation process without requiring manual intervention.

How does Litho ensure research agents receive the correct context from previous steps?

The orchestrator manages data dependencies through AgentDataConfig, where each agent declares required_sources and optional_sources such as DataSource::CODE_INSIGHTS or DataSource::ResearchResult. The orchestrator executes agents in a deterministic order that guarantees prerequisites complete first. Results are stored in MemoryScope::STUDIES_RESEARCH, allowing downstream agents to retrieve structured data via the context's memory system.

Can the research pipeline handle large codebases efficiently?

Yes, the pipeline incorporates parallelization specifically for intensive analysis tasks. The KeyModulesInsight agent in src/generator/research/agents/key_modules_insight.rs uses do_parallel_with_limit to analyze multiple domains simultaneously, respecting the max_parallels configuration setting. This prevents API rate limiting while maximizing throughput. Additionally, the conditional execution of DatabaseOverviewAnalyzer ensures SQL analysis only runs when relevant artifacts are detected, avoiding unnecessary processing.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →