How CodeWiki Generates System Architecture Diagrams from GitHub Repositories

CodeWiki automatically generates interactive system architecture diagrams from GitHub repositories by fetching repository metadata, processing it through a multi-step Gemini LLM workflow to produce Mermaid.js diagrams, and rendering them with clickable links back to source files.

CodeWiki is an open-source tool that transforms any public or private GitHub repository into a visual system architecture diagram. By leveraging Google's Gemini-2.5-pro model and Mermaid.js, it analyzes repository structure and source code to generate interactive diagrams that map components directly to their corresponding files.

The Three-Phase Architecture Generation Pipeline

CodeWiki builds system architecture diagrams through a coordinated pipeline that chains repository analysis, LLM processing, and interactive rendering.

Phase 1: Repository Metadata Collection

The backend fetches the repository's file tree and README content using the GitHub REST API. In api/services/github_service.py, the get_tree_data method requests the recursive tree of the default branch, while get_readme retrieves and Base64-decodes the README file.

Phase 2: LLM-Powered Analysis and Diagram Synthesis

The core intelligence runs in api/generate_diagram.py using a two-step Gemini workflow:

  1. Explanation and Mapping: Gemini-2.5-pro processes SYSTEM_SECOND_PROMPT to generate a high-level system explanation and map components to specific source files.
  2. Diagram Generation: The model then processes SYSTEM_THIRD_PROMPT to emit only valid Mermaid.js code, including click events for interactivity.

Phase 3: Post-Processing and Interactive Rendering

The raw Mermaid code undergoes sanitization via process_click_events and handle_mermaid_validation in generate_diagram.py. These functions convert relative paths to full GitHub URLs and fix common Mermaid syntax errors before streaming the result to the frontend.

Frontend Integration: Triggering Diagram Generation

The React frontend initiates diagram generation through the useDiagram hook located in frontend/hooks/useDiagram.ts. This hook sends a POST request to /api/diagram/generate with the repository owner, name, and optional personal access token.

// frontend/hooks/useDiagram.ts
const requestBody = { owner, repo, githubPat };
await fetch(`${baseUrl}/api/diagram/generate`, {
  method: 'POST',
  headers: { 'Content-Type': 'application/json' },
  body: JSON.stringify(requestBody),
});

The hook consumes the Server-Sent Events (SSE) stream, updating React state with explanation chunks, component mappings, and the final diagram code.

Backend Implementation: The Generation Workflow

The generate_diagram handler in api/generate_diagram.py orchestrates the entire workflow as an async generator that yields SSE messages. The handler implements the following sequence:

  1. Cache GitHub data: Calls get_cached_github_data to retrieve or store the repository tree and README.
  2. Stream explanation: Invokes gemini_service.generate with SYSTEM_SECOND_PROMPT to produce the textual system description.
  3. Map components: Uses the same prompt to generate component-to-file mappings.
  4. Generate diagram: Calls Gemini with SYSTEM_THIRD_PROMPT to receive Mermaid code.
  5. Post-process: Applies process_click_events to inject GitHub URLs and handle_mermaid_validation to sanitize syntax.
  6. Stream response: Yields SSE messages for each phase, culminating in the complete diagram.

Fetching Repository Data from GitHub

The GithubService class in api/services/github_service.py encapsulates all GitHub API interactions. The service handles authentication via optional personal access tokens for private repositories.

Key methods include:

  • get_tree_data: Requests the recursive tree of the default branch (main or master) and returns a newline-separated list of file paths.
  • get_readme: Fetches the README.md file, Base64-decodes the content, and returns the raw markdown.

Prompt Engineering for Architecture Extraction

The LLM behavior is controlled by system prompts defined in utils/prompts.py. These prompts guide Gemini-2.5-pro through the extraction and visualization process.

  • SYSTEM_SECOND_PROMPT: Instructs the model to analyze the repository explanation and file tree, then output a component-to-path mapping that identifies which files implement which system components.
  • SYSTEM_THIRD_PROMPT: Directs the model to generate only valid Mermaid.js diagram code, including click events for each component that link to the corresponding GitHub URLs, and to apply specific styling rules.

Streaming LLM Responses with Gemini

The GeminiService in api/services/gemini_service.py manages the async streaming client for Google's Generative AI API. The service initializes with the GEMINI_API_KEY environment variable and exposes a generate method that yields text chunks as they arrive.


# api/services/gemini_service.py

response = await self.client.aio.models.generate_content_stream(
    model=self.model,
    config=types.GenerateContentConfig(
        system_instruction=system_prompt,
        response_mime_type="application/json",
        max_output_tokens=12000,
    ),
    contents=user_message,
)

This streaming approach allows the frontend to display the explanation and diagram progressively rather than waiting for the entire LLM generation to complete.

Before finalizing the diagram, the backend processes the raw Mermaid code to add interactivity. The process_click_events function in api/generate_diagram.py rewrites every click directive to point to full GitHub URLs, distinguishing between files (blob) and directories (tree).

Additionally, handle_mermaid_validation sanitizes common syntax issues such as invalid arrows, stray percentage signs, escaped quotes, and redundant direction TD statements that could break the Mermaid renderer.

Caching Generated Diagrams for Performance

To avoid redundant LLM calls for previously analyzed repositories, CodeWiki implements a caching layer defined in utils/constants.py. The DIAGRAM_CACHE_DIR constant specifies the .cache/diagram_cache directory where generated diagrams are persisted.

The API exposes /api/diagram/cached endpoints (GET and POST) to read or write cached diagrams, enabling fast reloads without re-invoking the Gemini model.

Rendering Mermaid Diagrams in React

The frontend renders the final architecture diagram using the Mermaid component in frontend/components/Mermaid.tsx. This component initializes the Mermaid library with a Japanese-style theme and generous text limits, then renders the diagram via mermaid.render.

Features include:

  • Dark mode adjustments
  • Optional zoom functionality via svg-pan-zoom
  • Fullscreen modal view for large diagrams
  • Error display with original source code for debugging

Summary

  • CodeWiki generates system architecture diagrams by analyzing GitHub repositories through a three-phase pipeline: metadata collection, LLM processing, and interactive rendering.
  • The frontend triggers generation via the useDiagram hook, which streams Server-Sent Events from the /api/diagram/generate endpoint.
  • The backend fetches repository data using GithubService, then processes it through GeminiService with specialized prompts in utils/prompts.py to produce Mermaid.js diagrams.
  • Post-processing enriches diagrams with clickable GitHub links via process_click_events and validates syntax through handle_mermaid_validation.
  • Caching via DIAGRAM_CACHE_DIR prevents redundant LLM calls for previously analyzed repositories.
  • The React frontend renders diagrams using the Mermaid component with support for zoom, dark mode, and fullscreen viewing.

Frequently Asked Questions

How does CodeWiki handle private GitHub repositories?

CodeWiki supports private repositories by accepting an optional githubPat (Personal Access Token) parameter in the useDiagram hook and API requests. The GithubService class passes this token in the Authorization header when fetching the repository tree and README from the GitHub REST API, allowing access to private content while maintaining the same diagram generation workflow used for public repositories.

What LLM model does CodeWiki use for generating architecture diagrams?

CodeWiki uses Google's Gemini-2.5-pro model through the GeminiService class in api/services/gemini_service.py. The implementation uses the generate_content_stream method to stream responses progressively, allowing the frontend to display explanation text and diagram chunks as they are generated rather than waiting for the complete response.

Can I customize the styling of generated architecture diagrams?

While CodeWiki applies mandatory styling rules through SYSTEM_THIRD_PROMPT in utils/prompts.py to ensure diagram validity, the Mermaid React component in frontend/components/Mermaid.tsx initializes Mermaid with a Japanese-style theme and supports dark mode adjustments. The component also provides interactive features like zoom and fullscreen viewing, though direct customization of the generated Mermaid syntax would require modifying the backend prompt definitions.

How does CodeWiki prevent redundant processing of the same repository?

CodeWiki implements a caching mechanism defined in utils/constants.py using the DIAGRAM_CACHE_DIR constant (defaulting to .cache/diagram_cache). The API exposes /api/diagram/cached endpoints that check for existing diagrams before invoking the LLM workflow. When a diagram is generated, it is persisted to disk, enabling instant retrieval for subsequent requests without re-invoking Gemini or re-fetching GitHub data.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →