Multi-Step Diagram Generation Process in CodeWiki: A 4-Stage Pipeline Explained

The multi-step diagram generation process in CodeWiki follows a four-stage pipeline: retrieve cached GitHub data, generate a system design explanation via LLM, map components to specific files, and produce a validated Mermaid diagram with post-processing.

The quangdungluong/codewiki repository implements an intelligent architecture visualization system that transforms raw GitHub repositories into interactive Mermaid diagrams. Understanding this multi-step diagram generation process is essential for developers looking to extend the pipeline or debug generation issues. The workflow is orchestrated by the /api/diagram/generate endpoint in api/generate_diagram.py, which coordinates multiple LLM calls and post-processing steps to deliver production-ready architecture diagrams.

Overview of the 4-Stage Pipeline

The diagram generation endpoint executes a sequential pipeline where each stage feeds into the next. The process leverages Google's Gemini LLM for three distinct generation phases, with caching and post-processing layers to ensure accuracy and performance. Server-Sent Events (SSE) stream progress updates to the client in real-time, allowing the frontend to render intermediate explanations before the final diagram arrives.

Step 1: Retrieve and Cache Project Data

Caching Strategy with get_cached_github_data

The pipeline begins by fetching repository metadata through the get_cached_github_data helper function (lines 25-30 in api/generate_diagram.py). This function checks DIAGRAM_CACHE_DIR (defined in utils/constants.py) for existing data before making fresh API calls, significantly reducing latency for repeated requests against the same repository.

GitHub Data Sources

The GithubService class (located in api/services/github_service.py) retrieves three critical data points:

  • File tree: The complete repository structure to understand module organization
  • README content: Project documentation providing architectural context
  • Default branch: Reference for constructing accurate file URLs in later click handlers

This cached data package serves as the foundation for all subsequent LLM prompts.

Step 2: Generate System Design Explanation

LLM Prompting with SYSTEM_SECOND_PROMPT

The second stage initiates the first LLM call using SYSTEM_SECOND_PROMPT (defined in utils/prompts.py). The prompt instructs Gemini to analyze the file tree and README, then produce a detailed system-design explanation describing the architecture, data flow, and component relationships.

Streaming Response Handling

The generation occurs at lines 53-59 in api/generate_diagram.py, where the service streams chunks back to the client immediately. This allows the frontend to display the architectural explanation while the pipeline continues processing in the background. The SSE status explanation signals this stage's completion.

Step 3: Map Components to Files

Component-to-File Mapping Logic

Stage three executes a second LLM call (lines 64-70) that feeds the previously generated explanation back to Gemini with a specific instruction: map every architectural component to its corresponding file or directory path in the repository. This bridges the gap between abstract architectural concepts and concrete code locations.

Parsing Component Mapping Tags

The response is wrapped in <component_mapping> XML tags for reliable parsing. This structured output ensures the final diagram generation stage receives a clean dictionary correlating components like "Authentication Service" with specific paths like src/auth/.

Step 4: Generate and Post-Process Mermaid Diagram

Diagram Generation with SYSTEM_THIRD_PROMPT

The final LLM invocation (lines 72-84) uses SYSTEM_THIRD_PROMPT to synthesize the explanation and component mapping into valid Mermaid syntax. This prompt specifically instructs the model to generate flowchart or architecture diagram code that accurately represents the relationships identified in previous stages.

URL Rewriting with process_click_events

Before returning the diagram, the service executes process_click_events (lines 32-49) to rewrite click URLs. This function distinguishes between files and directories, appending /tree/ or /blob/ segments to ensure GitHub links resolve correctly when users interact with the generated diagram.

Syntax Validation with handle_mermaid_validation

The handle_mermaid_validation function (lines 52-75) performs final cleanup:

  • Removes invalid arrow syntax and stray % characters
  • Fixes escaped quotes that break Mermaid parsers
  • Strips redundant direction TD declarations
  • Validates overall diagram structure

The cleaned diagram is then streamed to the client with the complete status event.

Real-Time Streaming with Server-Sent Events

The endpoint communicates progress through SSE, allowing the frontend to render loading states and intermediate content. Each stage triggers a specific status update: started, explanation, mapping, diagram, and complete.

JavaScript client implementation:

async function generateDiagram(owner, repo, token) {
  const response = await fetch('/api/diagram/generate', {
    method: 'POST',
    headers: { 'Content-Type': 'application/json' },
    body: JSON.stringify({ owner, repo, token })
  });

  const reader = response.body.getReader();
  const decoder = new TextDecoder('utf-8');

  while (true) {
    const { done, value } = await reader.read();
    if (done) break;
    const chunk = decoder.decode(value);
    
    chunk.trim().split('\n').forEach(line => {
      if (line.startsWith('data:')) {
        const data = JSON.parse(line.slice(5));
        console.log('status →', data.status, data);
        if (data.status === 'complete') {
          renderMermaid(data.diagram);
        }
      }
    });
  }
}

Python client using httpx:

import httpx
import json

def get_diagram(owner, repo, token):
    url = "http://localhost:8000/api/diagram/generate"
    payload = {"owner": owner, "repo": repo, "token": token}
    
    with httpx.stream("POST", url, json=payload) as r:
        for line in r.iter_lines():
            if line.startswith(b"data:"):
                data = json.loads(line[5:])
                if data.get("status") == "complete":
                    print("Mermaid diagram:\n", data["diagram"])
                    break

Key Files and Architecture

Understanding the repository structure helps when extending the pipeline or debugging failures:

  • api/generate_diagram.py — Core orchestration endpoint implementing the four-stage pipeline, caching logic, and post-processing functions (process_click_events, handle_mermaid_validation).

  • utils/prompts.py — Contains SYSTEM_SECOND_PROMPT and SYSTEM_THIRD_PROMPT templates that guide LLM behavior for explanation and diagram generation.

  • api/services/gemini_service.py — Wrapper for the Gemini API handling all three LLM invocations with streaming support.

  • api/services/github_service.py — Retrieves repository metadata, file trees, and README content used as input for the generation process.

  • utils/constants.py — Defines DIAGRAM_CACHE_DIR for persisting cached repository data.

  • utils/logger.py — Provides structured logging for pipeline stages and debugging output.

Summary

The multi-step diagram generation process in CodeWiki transforms GitHub repositories into interactive Mermaid diagrams through a rigorous four-stage pipeline:

  • Data Retrieval: Caches GitHub file trees and README content to minimize API calls and latency.
  • Explanation Generation: Uses Gemini LLM with SYSTEM_SECOND_PROMPT to produce architectural explanations streamed in real-time.
  • Component Mapping: Correlates abstract architectural components to concrete file paths using structured XML tags.
  • Diagram Synthesis: Generates valid Mermaid code with SYSTEM_THIRD_PROMPT, then post-processes with process_click_events for URL rewriting and handle_mermaid_validation for syntax cleanup.

The entire workflow communicates via Server-Sent Events, allowing clients to render progress incrementally while maintaining clean separation between data fetching, LLM orchestration, and diagram validation.

Frequently Asked Questions

How does CodeWiki avoid repeated GitHub API calls for the same repository?

CodeWiki implements a caching layer through the get_cached_github_data function in api/generate_diagram.py. This helper checks DIAGRAM_CACHE_DIR (defined in utils/constants.py) for existing repository data before invoking the GithubService. By persisting file trees and README content locally, the multi-step diagram generation process eliminates redundant network requests and significantly reduces latency for subsequent diagram generations of the same repository.

What LLM prompts drive the diagram generation pipeline?

The pipeline utilizes two primary prompt templates defined in utils/prompts.py. SYSTEM_SECOND_PROMPT instructs Gemini to analyze the repository structure and produce a detailed system-design explanation. SYSTEM_THIRD_PROMPT then directs the model to synthesize that explanation—along with the component-to-file mapping—into valid Mermaid diagram syntax. These prompts are executed sequentially in api/generate_diagram.py at lines 53-59 and 72-84 respectively.

How does the system ensure generated Mermaid diagrams are valid and clickable?

Post-processing occurs in two specialized functions within api/generate_diagram.py. First, process_click_events (lines 32-49) rewrites URL paths in Mermaid click events to correctly distinguish between GitHub blob (file) and tree (directory) URLs. Second, handle_mermaid_validation (lines 52-75) cleans syntax errors including invalid arrows, stray percent signs, escaped quotes, and redundant direction declarations. This dual-layer validation ensures the final diagram renders correctly in Mermaid parsers while maintaining interactive navigation to source files.

Can I monitor the diagram generation progress in real-time?

Yes, the /api/diagram/generate endpoint implements Server-Sent Events (SSE) to stream status updates throughout the multi-step diagram generation process. The endpoint emits events with statuses including started, explanation, mapping, diagram, and complete. Clients can consume these streams using standard fetch APIs in JavaScript or httpx in Python, allowing frontend applications to display loading states, render intermediate explanations, and present the final Mermaid diagram the moment generation completes.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →