How code-review-graph Traces Execution Flows Through a Codebase: A Deep Dive
code-review-graph identifies, follows, and scores execution paths by detecting entry points, running breadth-first search traversal, and computing weighted criticality scores with in-memory caching.
The code-review-graph toolkit transforms static code analysis into navigable execution flows — ordered paths showing how programs actually run. The approach combines graph query detection with forward BFS traversal and multi-factor scoring, all powered by a SQLite-backed knowledge graph. This article explains the complete flow tracing pipeline as implemented in the tirth8205/code-review-graph repository.
Entry-point Detection: Finding Where Programs Start
Before tracing any paths, the system must identify entry points — functions capable of initiating execution. The detect_entry_points function in code_review_graph/flows.py (lines 65-78) applies three criteria:
| Criterion | Description |
|---|---|
| No incoming CALLS edges | true root nodes with no callers |
| Framework decorators | @app.get, @click.command, @KafkaListener, etc. |
| Conventional names | main, handler, lambda_handler, lifecycle hooks |
Test nodes are excluded by default. Pass include_tests=True to override.
from code_review_graph.flows import detect_entry_points
# Find all entry points in the loaded graph
entry_points = detect_entry_points(store, include_tests=False)
print(f"Found {len(entry_points)} entry points")
The function returns node IDs for each detected entry point, which serve as BFS starting positions.
Flow Tracing: Breadth-First Search Through the Graph
The trace_flows function (lines 285-317 in code_review_graph/flows.py) orchestrates forward traversal from every entry point. For each starting node, it delegates to the internal helper _trace_single_flow.
The tracing process:
- Loads FlowAdjacency from
GraphStore.load_flow_adjacency— an in-memory cache of outgoing CALLS edges - Expands nodes breadth-first up to
max_depth(default 15) - Records ordered node ID lists as
pathsequences - Filters out trivial single-node paths
- Computes criticality scores for remaining flows
from code_review_graph.flows import trace_flows
# Trace all flows with default depth limit
flows = trace_flows(store, max_depth=15, include_tests=False)
# Inspect a traced flow
for flow in flows:
print(f"{flow['name']}: depth={flow['depth']}, files={len(flow['files'])}")
The BFS avoids per-step database queries by using the pre-loaded adjacency structure.
Criticality Scoring: Ranking Flow Importance
Not all execution paths deserve equal attention. The compute_criticality function (lines 25-94 in code_review_graph/flows.py) calculates a 0.0-1.0 score across five weighted dimensions:
| Dimension | Weight | Measurement |
|---|---|---|
| File spread | 0.30 | count of distinct source files touched |
| External calls | 0.20 | calls to nodes absent from the graph |
| Security sensitivity | 0.25 | presence of security-related keywords from constants.py |
| Test-coverage gap | 0.15 | fraction of nodes lacking tests |
| Depth | 0.10 | BFS depth reached |
Higher scores indicate flows that are broader, riskier, or less tested — prime targets for code review attention.
# Access criticality scores on traced flows
critical_flows = sorted(flows, key=lambda f: f["criticality"], reverse=True)
print("Top 3 most critical flows:")
for flow in critical_flows[:3]:
print(f" {flow['name']}: {flow['criticality']:.3f}")
In-Memory Adjacency: Speeding Up Traversal
The FlowAdjacency class in code_review_graph/graph.py (lines 55-67) eliminates database round-trips during BFS. Built by GraphStore.load_flow_adjacency, it caches:
- Outgoing CALLS edges per node
- Test-coverage information
- Fast node lookups
This structure transforms the tracing operation from I/O-bound to memory-bound, essential for traversing large codebases efficiently.
Persistence: Storing Flows for Querying
Discovered flows are written to the SQLite knowledge graph via store_flows (lines 2-50 in code_review_graph/flows.py). The operation:
- Inserts flow records into the
flowstable - Populates
flow_membershipslinking flows to their constituent nodes - Executes within a single explicit transaction for consistency
from code_review_graph.flows import store_flows
# Persist all discovered flows
count = store_flows(store, flows)
print(f"Stored {count} flows in database")
Once persisted, flows become queryable through CLI commands and programmatic APIs.
Complete Working Example
This end-to-end example demonstrates the full flow tracing pipeline:
from code_review_graph.graph import GraphStore
from code_review_graph.flows import trace_flows, store_flows
# Initialize connection to knowledge graph
store = GraphStore(db_path=".code-review-graph/db.sqlite")
# Detect entry points, trace flows, compute criticality
flows = trace_flows(store, max_depth=15, include_tests=False)
# Persist for downstream analysis
stored = store_flows(store, flows)
# Report findings
print(f"Discovered and stored {stored} execution flows")
most_critical = max(flows, key=lambda f: f["criticality"])
print(f"Highest criticality: {most_critical['name']} ({most_critical['criticality']:.3f})")
For command-line usage:
# List all flows sorted by criticality
code-review-graph list_flows --sort-by criticality
# Find flows affected by a specific file change
code-review-graph get_affected_flows --file src/auth.py
Summary
- Entry-point detection in
detect_entry_pointsidentifies program starting points via graph topology, decorators, and naming conventions - Forward BFS tracing in
trace_flowswalks execution paths up to configurable depth limits - Criticality scoring weighs file spread, external calls, security keywords, test coverage, and depth
- FlowAdjacency caching eliminates database queries during traversal
- SQLite persistence in
store_flowsmakes flows queryable for impact analysis and review automation
Frequently Asked Questions
What makes a function an entry point in code-review-graph?
A function qualifies when it has no incoming CALLS edges, carries a framework decorator like @app.get or @click.command, or matches conventional names such as main or handler. The detect_entry_points function evaluates these criteria in code_review_graph/flows.py.
How does code-review-graph handle deeply nested call chains?
The BFS tracer respects a max_depth parameter (default 15) to bound exploration. Deeper paths are truncated, preventing unbounded traversal in recursive or highly connected codebases.
Can I include test functions in flow tracing?
Yes. Pass include_tests=True to trace_flows or detect_entry_points. By default, test nodes are excluded to focus on production execution paths.
Where are security keywords defined for criticality scoring?
Security-related terms are maintained in code_review_graph/constants.py and referenced by compute_criticality. These keywords flag flows that warrant extra scrutiny during security reviews.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →