How code-review-graph Detects Execution Flows from Entry Points: A Deep Dive into the Source Code
code-review-graph detects execution flows by identifying entry-point nodes through heuristics like uncalled functions and framework patterns, then tracing reachable calls via depth-first traversal of the stored code graph.
This open-source tool, hosted at tirth8205/code-review-graph, automates the discovery of how code actually runs in a project. Understanding its detection mechanism helps developers leverage the tool for impact analysis, visualization, and automated code reviews.
Entry-Point Detection: The Starting Point for Flow Analysis
The core detection logic resides in detect_entry_points() within [code_review_graph/flows.py (lines 165-215)](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/flows.py#L165-L215). This function walks the stored graph and applies four distinct heuristics to identify where execution can begin.
Stand-Alone Roots (Functions with No Callers)
The simplest heuristic targets orphan functions — nodes that exist in the call graph but have no incoming edges. These represent potential manual invocation points or scripts that execute directly.
Framework-Specific Pattern Matching
The tool recognizes decorator and naming conventions from popular frameworks:
- Flask:
@app.route - FastAPI: endpoint decorators
- Django:
@receiver - Celery:
@task - pytest: fixtures
- Spring:
@Scheduled - Lifespan:
@lifespan
This pattern matching enables accurate detection without requiring runtime execution.
Explicit Configuration and Test Inclusion
Users can manually mark files as entry points through configuration files. Additionally, test files are optionally included when include_tests=True, enabling analysis of test execution paths alongside production code.
The function returns a list of GraphNode objects, each representing a validated entry point with its qualified name, file location, and metadata.
Flow Tracing: Walking the Call Graph from Entry Points
Once entry points are identified, trace_flows() (lines 285-313) performs the actual flow construction through depth-first traversal.
Building Flow Records
For each entry point, the tracer constructs a comprehensive flow record containing:
- Qualified name: The full reference to the entry function
- Depth: Maximum call nesting level reached
- Node count: Total unique functions visited
- File count: Distinct source files touched
- Ordered steps: The complete call sequence with source locations
The traversal respects language-specific resolvers (Python, Java, Rust) that have pre-populated the graph with caller/callee relationships during the initial build phase.
Persisting Flow Metadata
The store_flows() function (lines 272-294) persists these records into the internal SQLite store. Each flow is linked to its entry_point_id, enabling efficient lookup and filtering in subsequent operations.
MCP Tool Wrappers: Exposing Flows Through the API
The public interface consists of thin wrappers around the core functions, defined in [code_review_graph/tools/flows_tools.py (lines 53-65)](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/tools/flows_tools.py#L53-L65).
List All Detected Flows
from code_review_graph.tools.flows_tools import list_flows
result = list_flows(repo_root="/path/to/repo")
print(result["flows"]) # each entry contains name, entry_point, depth, etc.
Retrieve a Specific Flow's Call Sequence
from code_review_graph.tools.flows_tools import get_flow
flow = get_flow(flow_name="handle_request", repo_root="/path/to/repo")
print(flow["flow"]["steps"]) # ordered list of call steps with source locations
Identify Affected Flows from Changed Files
from code_review_graph.tools.flows_tools import get_affected_flows
affected = get_affected_flows(changed_files=["auth.py"], repo_root="/path/to/repo")
print([f["name"] for f in affected["affected_flows"]])
Supporting Components and Architecture
| File | Purpose |
|---|---|
code_review_graph/flows.py |
Core logic for detect_entry_points, trace_flows, and store_flows |
code_review_graph/tools/flows_tools.py |
MCP tool wrappers exposing flows via CLI and API |
code_review_graph/refactor.py |
Contains _is_entry_point helper mirroring detection heuristics for refactoring operations |
code_review_graph/tools/build.py |
Persists flow records into SQLite during build steps |
The refactor.py module deserves special mention — its _is_entry_point function duplicates the detection heuristics, ensuring consistent entry-point recognition across analysis and refactoring workflows.
Summary
detect_entry_points()inflows.pyapplies heuristics (uncalled functions, framework patterns, explicit marks) to find execution starting pointstrace_flows()performs depth-first traversal to build complete call sequences from each entry point- Flow records capture depth, node count, file count, and ordered steps for downstream analysis
- MCP wrappers in
flows_tools.pyexpose functionality throughlist_flows,get_flow, andget_affected_flows - SQLite persistence enables fast querying of flow metadata during code review operations
Frequently Asked Questions
How does code-review-graph handle multiple programming languages?
The tool uses language-specific resolvers that populate the graph with caller/callee edges during the build phase. These resolvers parse Python, Java, Rust, and other supported languages into a unified graph structure. The flow detection logic operates on this normalized representation, making it language-agnostic at the traversal level.
Can I customize which files are treated as entry points?
Yes. Beyond automatic heuristics, you can explicitly mark files as entry points through configuration. The include_tests parameter also controls whether test files participate in flow detection, useful for analyzing test coverage impact or excluding test-only paths from production flow analysis.
What performance characteristics does flow tracing have?
Flow tracing uses depth-first traversal with cycle detection to prevent infinite loops. The operation's complexity scales with the number of entry points multiplied by call graph depth. Results are cached in SQLite, making subsequent queries (get_flow, get_affected_flows) near-instantaneous after initial analysis.
How accurate is the affected-flow detection for code changes?
The get_affected_flows tool identifies flows where any step references a changed file. This is over-approximation by design — it catches all potentially impacted execution paths without requiring expensive inter-procedural analysis. Developers receive a conservative superset of affected flows, ensuring no critical paths are missed during review.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →