How the Review Context Generation Flow Works in code-review-graph
The review context generation flow in code-review-graph operates in three stages: detecting changed files, constructing a focused impact sub-graph, and assembling source snippets into a token-efficient LLM prompt.
The code-review-graph project implements a token-efficient pipeline for generating review contexts. This system ensures that large language models receive only the relevant portions of a codebase when performing code reviews, dramatically reducing token consumption while preserving critical context. The entire flow is orchestrated through the public API get_review_context_tool defined in code_review_graph/main.py.
Detect Changed Files
The review context generation flow begins with change detection. The detect_changes_tool function, implemented in code_review_graph/tools/analysis_tools.py, analyzes the current Git commit or pull request. It collects the complete set of files that were added, removed, or modified. Each file is annotated with exact line-range changes to establish precise boundaries for downstream processing.
This stage establishes the foundation for all subsequent context extraction. Without accurate change detection, the sub-graph construction would include irrelevant code paths or miss critical dependencies.
Construct a Focused Sub-graph
With the list of changed files identified, the flow proceeds to impact analysis. The get_review_context function calls the graph-query engine in code_review_graph/graph.py to retrieve the minimal set of nodes and edges reachable from the changed files.
This impact sub-graph contains:
- Directly affected symbols
- Callers and callees of changed functions
- Cross-language dependencies (for example, Python ↔ JavaScript)
The sub-graph is then condensed into a hierarchical structure suitable for textual rendering. This condensation step is critical for token efficiency—it transforms raw graph data into a compact summary that retains structural relationships without verbose serialization.
Gather Source Snippets and Assemble the Prompt
The final stage transforms the sub-graph into concrete review material. For each node in the sub-graph, get_review_context extracts relevant source excerpts surrounding the changed lines. These snippets are trimmed to fit within the LLM's context window constraints.
Each snippet receives metadata labeling:
- File path
- Line numbers
- Role description (for example, "function
process_datacallsvalidate")
The review prompt combines four components:
- High-level change description
- Condensed sub-graph summary
- Curated source snippets
- Fixed instruction block from
code_review_graph/prompts.py
The result is returned as a ReviewContext object ready for LLM consumption.
Public Entry Point and Error Handling
The get_review_context_tool function wraps the core routine with provenance tracking and error handling. This design makes the flow safe to invoke from multiple interfaces:
- The CLI command
crg review - Programmatic usage via the MCP toolset
- Direct Python import
Python API Usage
from code_review_graph.main import get_review_context_tool
# Generate a review context for the current PR, using the default base branch.
review_ctx = get_review_context_tool(base="main")
print(review_ctx.prompt) # Full LLM-ready prompt
print(review_ctx.snippets) # Mapping of file → trimmed source excerpts
CLI Usage
# Token-optimised review context
crg review --base main > review_prompt.txt
Manual Integration
# Manual integration in a custom script
changes = detect_changes_tool()
subgraph = build_subgraph(changes)
ctx = get_review_context(base="main", changes=changes, subgraph=subgraph)
llm_response = call_llm(ctx.prompt)
Pipeline Architecture
The complete flow can be visualized as:
detect_changes_tool → sub-graph extraction → source-snippet collection
│ │ │
└───────► get_review_context_tool ◄───────────────────┘
This architecture enforces a strict separation of concerns while maintaining composability. Each stage consumes the output of the previous stage, enabling flexible integration into custom workflows.
Key Source Files
The review context generation flow spans five essential files in the repository:
code_review_graph/tools/review.py— Core implementation ofget_review_contextcode_review_graph/main.py— Public wrapperget_review_context_toolcode_review_graph/graph.py— Graph data structures and sub-graph extraction logiccode_review_graph/prompts.py— Fixed instruction block prepended to generated promptsskills/review-pr/SKILL.md— Example usage pattern for PR-review workflows
Summary
- The review context generation flow uses three coordinated stages to minimize token usage while preserving review accuracy
- Change detection in
analysis_tools.pyestablishes precise boundaries for modified code - Sub-graph extraction in
graph.pyidentifies only relevant dependent symbols across language boundaries - Source snippet assembly applies windowing and metadata labeling to prepare LLM-ready content
- The
get_review_context_toolwrapper provides safe, tracked access from CLI, API, or programmatic contexts
Frequently Asked Questions
What makes the review context generation flow token-efficient?
The flow achieves token efficiency through selective inclusion. Instead of passing entire files to the LLM, it constructs an impact sub-graph containing only symbols reachable from changed files. Source snippets are windowed around change points and condensed hierarchically. According to the code-review-graph source code, this typically reduces context size by 80-95% compared to full-file approaches.
How does code-review-graph handle cross-language dependencies?
The graph engine in code_review_graph/graph.py maintains a unified symbol graph across multiple languages. When building the impact sub-graph, it follows edges regardless of source language—enabling detection of JavaScript callers when Python functions change. This is critical for modern polyglot codebases where review accuracy depends on understanding full call chains.
Can I customize the prompt template used in review context generation?
The fixed instruction block lives in code_review_graph/prompts.py and is currently prepended to all generated prompts. For template customization, you can modify this file directly or post-process the ReviewContext.prompt field after generation. The project structure separates prompt construction from core logic, enabling straightforward extension.
What is the difference between get_review_context and get_review_context_tool?
get_review_context in code_review_graph/tools/review.py implements the core three-stage pipeline. get_review_context_tool in code_review_graph/main.py adds a provenance-tracking wrapper with standardized error handling and MCP tool compatibility. For most use cases, import from main.py; for custom orchestration requiring direct control, use the lower-level function from review.py.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →