How code-review-graph Detects Execution Flows from Entry Points with Criticality Weighting
Execution flow detection in code-review-graph identifies entry points via framework decorators and naming conventions, traces forward paths using breadth-first search, and quantifies risk through a weighted criticality algorithm that prioritizes high-impact code paths.
The code-review-graph open-source project implements static analysis to map how code actually runs by discovering execution flows from entry points with criticality weighting. This approach enables automated prioritization of review efforts toward the most risky and impactful paths in your codebase.
Locating Entry Points in the Call Graph
The detection process begins in code_review_graph/flows.py where detect_entry_points() (lines 65-78) scans all Function and Test nodes to identify where execution originates. A node qualifies as an entry point if it meets any of three distinct criteria.
Framework Decorator Detection
The system recognizes modern web frameworks and CLI tools through decorator analysis. The helper _has_framework_decorator_() (lines 30-41 in flows.py) identifies markers such as @app.get, @click.command, @receiver, and similar framework-specific annotations that indicate externally callable functions.
Conventional Naming Patterns
Beyond decorators, _matches_entry_name_() (lines 44-52) applies regex patterns to function names, catching conventional signatures like main, test_*, handler, lambda_handler, and Android lifecycle methods. This ensures the graph captures entry points even in codebases without explicit framework metadata.
Root Function Identification
Pure graph topology also defines entry points: any node with zero incoming CALLS edges represents a true root that initiates execution chains. By default, detect_entry_points() excludes test functions (include_tests=False) unless explicitly requested, filtering noise from the analysis.
Tracing Execution Paths with BFS
Once entry points are established, the system traces execution forward to build complete flow maps.
Configurable Traversal Depth
For each detected entry point, _trace_single_flow() (lines 23-63) executes a breadth-first search following calls_out edges. The traversal respects a configurable max_depth parameter (defaulting to 15 levels) to prevent infinite loops in recursive code while capturing meaningful call chains. Trivial single-node flows are automatically discarded.
Flow Data Collection
During BFS traversal, the algorithm collects:
- Node IDs and qualified names
- Distinct file paths touched by the flow
- BFS depth statistics
- External call indicators
The orchestrating function trace_flows() (lines 86-101) coordinates detection and traversal, returning flows sorted by their computed criticality scores in descending order.
Computing Criticality with Weighted Risk Factors
The compute_criticality() function (lines 25-94 in flows.py) transforms raw flow statistics into a normalized risk score between 0 and 1. The algorithm aggregates five specific risk dimensions with predetermined weights.
- File spread (weight 0.30): Measures the count of distinct source files the flow traverses. Greater file dispersion indicates broader architectural impact.
- External calls (weight 0.20): Tracks invocations targeting nodes absent from the graph, representing integration points with external services or libraries.
- Security sensitivity (weight 0.25): Scans node names and qualified names against security keywords defined in
code_review_graph/constants.py(e.g.,auth,secret,token). - Test-coverage gap (weight 0.15): Calculates the fraction of flow nodes lacking test coverage, highlighting untested execution paths.
- Depth (weight 0.10): Considers the BFS depth of the flow, with deeper call stacks receiving moderate risk elevation.
Each factor normalizes to [0, 1] before aggregation. The final sum clamps to [0, 1] and rounds to four decimal places, ensuring consistent precision for comparison and sorting.
Persisting Flow Data
The store_flows() function manages persistence by clearing previous flow data and writing newly discovered flows to SQLite tables flows and flow_memberships. For incremental analysis after code changes, incremental_trace_flows() (lines 60-74) selectively recomputes only flows affected by modified files, optimizing performance in large repositories.
Practical Implementation Examples
To analyze a codebase and retrieve high-criticality flows:
from code_review_graph.graph import GraphStore
from code_review_graph.flows import trace_flows
# Load existing graph store
store = GraphStore(path="myrepo.db")
# Detect and trace all flows, sorted by criticality
flows = trace_flows(store, max_depth=15, include_tests=False)
# Display top 5 most critical execution paths
for flow in flows[:5]:
print(
f"Flow {flow['name']!r} (criticality={flow['criticality']:.2f}) "
f"touches {flow['file_count']} files at depth {flow['depth']}"
)
For incremental updates after specific file changes:
from code_review_graph.flows import incremental_trace_flows
# Re-trace only flows affected by recent changes
changed_files = ["src/api/users.py", "src/auth/jwt.py"]
updated_count = incremental_trace_flows(store, changed_files)
print(f"Re-computed {updated_count} affected flows with updated criticality scores")
These APIs power the CLI command code-review-graph list_flows, surfacing ranked execution paths for manual review or automated gatekeeping.
Summary
- Entry point detection combines graph topology (root nodes), framework decorator analysis, and naming convention matching in
detect_entry_points(). - Flow tracing employs BFS traversal via
_trace_single_flow()with configurable depth limits to map executable paths from each entry point. - Criticality scoring applies weighted factors (file spread 0.30, security sensitivity 0.25, external calls 0.20, test gaps 0.15, depth 0.10) to quantify risk.
- Persistence uses SQLite tables
flowsandflow_memberships, withincremental_trace_flows()supporting efficient updates. - Ranking sorts flows by criticality descending, enabling reviewers to focus on the highest-risk execution paths first.
Frequently Asked Questions
What criteria define an entry point in code-review-graph?
An entry point is any function node that has no incoming CALLS edges, carries a recognized framework decorator (such as @app.get or @click.command), or matches conventional naming patterns like main, handler, or test_*. The detect_entry_points() function in code_review_graph/flows.py evaluates these criteria against the graph structure.
How does the criticality weighting system prioritize flows?
The system assigns higher scores to flows exhibiting greater architectural spread (30% weight), security-sensitive identifiers (25% weight), and external dependencies (20% weight). Test coverage gaps contribute 15%, while call depth adds 10%. This weighting surfaces flows that touch many files, handle authentication logic, or call external services as highest priority for review.
Can I customize the depth limit when tracing execution flows?
Yes. The trace_flows() function accepts a max_depth parameter that defaults to 15 levels. You can increase this for deep call stacks or decrease it for faster analysis of shallow execution paths. The BFS traversal in _trace_single_flow() respects this boundary to balance comprehensiveness against performance.
How does incremental tracing improve performance in large codebases?
The incremental_trace_flows() function accepts a list of changed file paths and recomputes criticality only for flows containing nodes from those files. This avoids full graph recomputation during iterative development, efficiently updating the flows table while preserving accurate criticality scores for affected execution paths.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →