How code-review-graph Generates Interactive Visualizations of Code Relationships
code-review-graph transforms a persisted SQLite knowledge graph into a self-contained, D3.js-powered HTML page with force-directed layouts, collapsible nodes, and full keyboard navigation — all working offline.
The open-source tool tirth8205/code-review-graph bridges the gap between static code analysis and dynamic exploration. After parsing a repository into a graph database of functions, classes, imports, and review flows, it renders an interactive visualization that developers can explore in any modern browser. The entire pipeline from raw graph data to clickable HTML runs through a single Python module: code_review_graph/visualization.py.
Exporting Graph Data from the SQLite Store
The visualization pipeline begins with export_graph_data() in code_review_graph/visualization.py (lines 172-225). This function walks the GraphStore and serializes every node and edge into plain dictionaries suitable for JSON encoding.
What gets exported:
- Nodes – entities with
id,kind(function, class, file, module),name,file_path,line, and optional community assignments - Edges – relationships with
source,target,kind(calls, imports, references, reviews), and resolved full identifiers for short targets - Flows – optional code review flow traces (reviewer → file → function chains)
- Communities – Louvain clustering results for modular decomposition
- Stats – degree distributions, density metrics, and coverage summaries
The function handles edge target resolution to ensure that unqualified names (like helper() in from .utils import helper) map to their fully qualified counterparts. This produces a single serializable payload:
{
"nodes": [...],
"edges": [...],
"stats": {...},
"flows": [...],
"communities": [...]
}
Choosing the Right Rendering Mode
Before generating HTML, code-review-graph must decide how much detail to render. The _resolve_auto_mode helper (lines 443-456) implements intelligent fallback based on graph size.
Default thresholds:
DEFAULT_MAX_FULL_NODES= 3000DEFAULT_MAX_FULL_EDGES= 9000 (3× node count)
Mode selection logic:
| Condition | Selected Mode | Result |
|---|---|---|
| Nodes ≤ 3000 AND Edges ≤ 9000 | full |
Complete graph with all individuals nodes |
| Community data exists | community |
Super-nodes representing detected communities |
| Otherwise | file |
Aggregation to file-level nodes only |
You can override this with the CLI flag:
code-review-graph visualize --mode community # force community view
code-review-graph visualize --mode full # attempt full graph regardless of size
The aggregation helpers _aggregate_community and _aggregate_file (lines 368-506) collapse the full graph while preserving drill-down capability. Community aggregation groups nodes by their Louvain-assigned community ID, computing internal edge density and cross-community links. File aggregation rolls function and class nodes up to their containing files, simplifying large codebases into manageable file-dependency networks.
Building the Interactive HTML Page
The generate_html() function (lines 58-70) orchestrates final output generation through three substages: template selection, asset injection, and D3 embedding.
Template Selection
Based on the resolved mode, the function chooses between:
_HTML_TEMPLATE– full graph with individual nodes_AGGREGATED_HTML_TEMPLATE– community or file super-nodes with expansion capability
D3 Asset Handling
The _d3_script_tags() function (lines 41-54) produces script tags with Subresource Integrity (SRI) pinning. _write_d3_asset() optionally copies the vendored code_review_graph/assets/d3.v7.min.js beside the output HTML. When --serve is used, this local copy enables fully offline operation. If the local copy is missing or checksum-invalid, the generator falls back to CDN with SRI verification.
Embedded Force-Directed Simulation
The generated HTML contains approximately 500 lines of inline JavaScript (lines ~8640-9100) that creates a live D3 force simulation:
// Core simulation setup (simplified from template)
const simulation = d3.forceSimulation(graph.nodes)
.force("link", d3.forceLink(graph.edges).id(d => d.id).distance(100))
.force("charge", d3.forceManyBody().strength(-300))
.force("collision", d3.forceCollide().radius(d => degreeRadius(d) + 5))
.force("center", d3.forceCenter(width / 2, height / 2));
Visual encoding strategy:
- Node radius – Scaled by degree centrality via
degreeRadius()function - Node color – Mapped from
KIND_COLORdictionary (functions, classes, files, modules each have distinct hues) - Node shape – SVG symbols from
KIND_SHAPE(circles, squares, diamonds) - Edge style –
EDGE_COLORper relationship type with arrowhead markers - Community coloring – Optional
communityColorScalewhen community data is present
User Interaction Features
The visualization supports multiple input modalities for code exploration:
Mouse interactions:
- Hover –
showTooltip()displays node details including kind, file path, line number, and community membership - Click –
toggleCollapse()expands/collapses file nodes to show/hide contained functions and classes;showDetailPanel()opens a persistent info panel for any node - Drag – D3 drag behavior pins nodes to new positions, with physics continuing around fixed points
Keyboard navigation:
- Arrow keys – Navigate between connected nodes
/– Focus search box for quick node finding?– Toggle help overlay with all shortcuts
Control panels:
- Legend buttons – Toggle visibility per edge kind via
hiddenEdgeKindsSet - Filter panel – Show/hide node kinds via
hiddenNodeKindsSet - Community toggle – Recolor nodes by community assignment or revert to kind-based coloring
CLI Usage and Built-in Server
The visualize sub-command in code_review_graph/cli.py (lines 891-910) wires everything together:
# Generate once, open file directly
code-review-graph visualize --output graph.html
# Build and serve with local D3 asset
code-review-graph visualize --serve
# → Starting server at http://127.0.0.1:8765/
# → Press Ctrl+C to stop
The --serve flag starts a minimal HTTP server that serves both the generated HTML and the local D3 bundle, ensuring no external network requests are needed.
Programmatic Integration
Embed visualization generation in Python workflows:
from code_review_graph.graph import GraphStore
from code_review_graph.visualization import generate_html
# Open existing analysis database
store = GraphStore.from_path("./myrepo.graph.db")
# Auto-detect appropriate mode
html_path = generate_html(
store,
output_path="analysis.html",
mode="auto",
max_full_nodes=5000, # raise threshold for larger displays
)
# Or force specific aggregation
community_html = generate_html(
store,
output_path="modules.html",
mode="community",
)
The render_pr_comment.py script in the repository demonstrates reusing generate_html() for CI/CD integrations, producing static HTML embeddable in pull request descriptions.
Summary
export_graph_data()serializes the SQLiteGraphStoreinto JSON-ready dictionaries with resolved references and optional flow/community annotations- Auto-mode resolution switches from full graphs to community or file aggregation when exceeding 3000 nodes or 9000 edges, protecting browser performance
generate_html()assembles self-contained HTML with vendored or CDN D3, substituting graph data into embedded templates- Inline D3 simulation provides force-directed layout, degree-scaled nodes, per-kind styling, and rich interaction hooks
- Full offline capability via
--servewith local D3 asset, requiring zero external dependencies at runtime
Frequently Asked Questions
What file size or node count limits should I expect?
The default thresholds are 3000 nodes and 9000 edges for full-graph mode. Beyond this, code-review-graph automatically switches to community or file aggregation. These limits are configurable via max_full_nodes and max_full_edges parameters. In practice, most repositories under 100k lines of code render fully interactive, while monorepos benefit from the aggregation modes.
Can I customize the colors or visual styling?
The visualization uses hardcoded KIND_COLOR, EDGE_COLOR, and KIND_SHAPE dictionaries in the embedded D3 code (around line 8650). For custom styling, you would need to modify code_review_graph/visualization.py and regenerate. The repository does not currently expose runtime theme configuration, though the HTML output is self-contained and can be hand-edited post-generation.
Does the visualization work without internet access?
Yes. When using --serve or when the vendored d3.v7.min.js is present alongside the output HTML, the visualization functions completely offline. The _write_d3_asset() function ensures the local D3 copy is available, and _d3_script_tags() generates SRI-pinned script tags that validate integrity without external requests.
How do community aggregations preserve detail for drill-down?
The _aggregate_community function computes super-node metadata including member lists, internal edge counts, and cross-community connections. When you click a community node in the aggregated view, the detail panel lists all contained entities with their individual statistics. For file aggregation, the toggleCollapse() function in the D3 code dynamically injects child nodes into the simulation, allowing progressive expansion without reloading the page.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →