How code-review-graph Generates Interactive Visualizations of Code Relationships

code-review-graph transforms a persisted SQLite knowledge graph into a self-contained, D3.js-powered HTML page with force-directed layouts, collapsible nodes, and full keyboard navigation — all working offline.

The open-source tool tirth8205/code-review-graph bridges the gap between static code analysis and dynamic exploration. After parsing a repository into a graph database of functions, classes, imports, and review flows, it renders an interactive visualization that developers can explore in any modern browser. The entire pipeline from raw graph data to clickable HTML runs through a single Python module: code_review_graph/visualization.py.

Exporting Graph Data from the SQLite Store

The visualization pipeline begins with export_graph_data() in code_review_graph/visualization.py (lines 172-225). This function walks the GraphStore and serializes every node and edge into plain dictionaries suitable for JSON encoding.

What gets exported:

  • Nodes – entities with id, kind (function, class, file, module), name, file_path, line, and optional community assignments
  • Edges – relationships with source, target, kind (calls, imports, references, reviews), and resolved full identifiers for short targets
  • Flows – optional code review flow traces (reviewer → file → function chains)
  • Communities – Louvain clustering results for modular decomposition
  • Stats – degree distributions, density metrics, and coverage summaries

The function handles edge target resolution to ensure that unqualified names (like helper() in from .utils import helper) map to their fully qualified counterparts. This produces a single serializable payload:

{
  "nodes": [...],
  "edges": [...],
  "stats": {...},
  "flows": [...],
  "communities": [...]
}

Choosing the Right Rendering Mode

Before generating HTML, code-review-graph must decide how much detail to render. The _resolve_auto_mode helper (lines 443-456) implements intelligent fallback based on graph size.

Default thresholds:

  • DEFAULT_MAX_FULL_NODES = 3000
  • DEFAULT_MAX_FULL_EDGES = 9000 (3× node count)

Mode selection logic:

Condition Selected Mode Result
Nodes ≤ 3000 AND Edges ≤ 9000 full Complete graph with all individuals nodes
Community data exists community Super-nodes representing detected communities
Otherwise file Aggregation to file-level nodes only

You can override this with the CLI flag:

code-review-graph visualize --mode community   # force community view

code-review-graph visualize --mode full        # attempt full graph regardless of size

The aggregation helpers _aggregate_community and _aggregate_file (lines 368-506) collapse the full graph while preserving drill-down capability. Community aggregation groups nodes by their Louvain-assigned community ID, computing internal edge density and cross-community links. File aggregation rolls function and class nodes up to their containing files, simplifying large codebases into manageable file-dependency networks.

Building the Interactive HTML Page

The generate_html() function (lines 58-70) orchestrates final output generation through three substages: template selection, asset injection, and D3 embedding.

Template Selection

Based on the resolved mode, the function chooses between:

  • _HTML_TEMPLATE – full graph with individual nodes
  • _AGGREGATED_HTML_TEMPLATE – community or file super-nodes with expansion capability

D3 Asset Handling

The _d3_script_tags() function (lines 41-54) produces script tags with Subresource Integrity (SRI) pinning. _write_d3_asset() optionally copies the vendored code_review_graph/assets/d3.v7.min.js beside the output HTML. When --serve is used, this local copy enables fully offline operation. If the local copy is missing or checksum-invalid, the generator falls back to CDN with SRI verification.

Embedded Force-Directed Simulation

The generated HTML contains approximately 500 lines of inline JavaScript (lines ~8640-9100) that creates a live D3 force simulation:

// Core simulation setup (simplified from template)
const simulation = d3.forceSimulation(graph.nodes)
  .force("link", d3.forceLink(graph.edges).id(d => d.id).distance(100))
  .force("charge", d3.forceManyBody().strength(-300))
  .force("collision", d3.forceCollide().radius(d => degreeRadius(d) + 5))
  .force("center", d3.forceCenter(width / 2, height / 2));

Visual encoding strategy:

  • Node radius – Scaled by degree centrality via degreeRadius() function
  • Node color – Mapped from KIND_COLOR dictionary (functions, classes, files, modules each have distinct hues)
  • Node shape – SVG symbols from KIND_SHAPE (circles, squares, diamonds)
  • Edge style – EDGE_COLOR per relationship type with arrowhead markers
  • Community coloring – Optional communityColorScale when community data is present

User Interaction Features

The visualization supports multiple input modalities for code exploration:

Mouse interactions:

  • Hover – showTooltip() displays node details including kind, file path, line number, and community membership
  • Click – toggleCollapse() expands/collapses file nodes to show/hide contained functions and classes; showDetailPanel() opens a persistent info panel for any node
  • Drag – D3 drag behavior pins nodes to new positions, with physics continuing around fixed points

Keyboard navigation:

  • Arrow keys – Navigate between connected nodes
  • / – Focus search box for quick node finding
  • ? – Toggle help overlay with all shortcuts

Control panels:

  • Legend buttons – Toggle visibility per edge kind via hiddenEdgeKinds Set
  • Filter panel – Show/hide node kinds via hiddenNodeKinds Set
  • Community toggle – Recolor nodes by community assignment or revert to kind-based coloring

CLI Usage and Built-in Server

The visualize sub-command in code_review_graph/cli.py (lines 891-910) wires everything together:


# Generate once, open file directly

code-review-graph visualize --output graph.html

# Build and serve with local D3 asset

code-review-graph visualize --serve

# → Starting server at http://127.0.0.1:8765/

# → Press Ctrl+C to stop

The --serve flag starts a minimal HTTP server that serves both the generated HTML and the local D3 bundle, ensuring no external network requests are needed.

Programmatic Integration

Embed visualization generation in Python workflows:

from code_review_graph.graph import GraphStore
from code_review_graph.visualization import generate_html

# Open existing analysis database

store = GraphStore.from_path("./myrepo.graph.db")

# Auto-detect appropriate mode

html_path = generate_html(
    store,
    output_path="analysis.html",
    mode="auto",
    max_full_nodes=5000,  # raise threshold for larger displays

)

# Or force specific aggregation

community_html = generate_html(
    store,
    output_path="modules.html",
    mode="community",
)

The render_pr_comment.py script in the repository demonstrates reusing generate_html() for CI/CD integrations, producing static HTML embeddable in pull request descriptions.

Summary

  • export_graph_data() serializes the SQLite GraphStore into JSON-ready dictionaries with resolved references and optional flow/community annotations
  • Auto-mode resolution switches from full graphs to community or file aggregation when exceeding 3000 nodes or 9000 edges, protecting browser performance
  • generate_html() assembles self-contained HTML with vendored or CDN D3, substituting graph data into embedded templates
  • Inline D3 simulation provides force-directed layout, degree-scaled nodes, per-kind styling, and rich interaction hooks
  • Full offline capability via --serve with local D3 asset, requiring zero external dependencies at runtime

Frequently Asked Questions

What file size or node count limits should I expect?

The default thresholds are 3000 nodes and 9000 edges for full-graph mode. Beyond this, code-review-graph automatically switches to community or file aggregation. These limits are configurable via max_full_nodes and max_full_edges parameters. In practice, most repositories under 100k lines of code render fully interactive, while monorepos benefit from the aggregation modes.

Can I customize the colors or visual styling?

The visualization uses hardcoded KIND_COLOR, EDGE_COLOR, and KIND_SHAPE dictionaries in the embedded D3 code (around line 8650). For custom styling, you would need to modify code_review_graph/visualization.py and regenerate. The repository does not currently expose runtime theme configuration, though the HTML output is self-contained and can be hand-edited post-generation.

Does the visualization work without internet access?

Yes. When using --serve or when the vendored d3.v7.min.js is present alongside the output HTML, the visualization functions completely offline. The _write_d3_asset() function ensures the local D3 copy is available, and _d3_script_tags() generates SRI-pinned script tags that validate integrity without external requests.

How do community aggregations preserve detail for drill-down?

The _aggregate_community function computes super-node metadata including member lists, internal edge counts, and cross-community connections. When you click a community node in the aggregated view, the detail panel lists all contained entities with their individual statistics. For file aggregation, the toggleCollapse() function in the D3 code dynamically injects child nodes into the simulation, allowing progressive expansion without reloading the page.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →