# How the Review Context Generation Flow Works in code-review-graph

> Understand the review context generation flow in code-review-graph. It detects files, builds impact sub-graphs, and assembles snippets for efficient LLM prompts. Learn more.

- Repository: [Tirth Kanani/code-review-graph](https://github.com/tirth8205/code-review-graph)
- Tags: deep-dive
- Published: 2026-08-11

---

**The review context generation flow in code-review-graph operates in three stages: detecting changed files, constructing a focused impact sub-graph, and assembling source snippets into a token-efficient LLM prompt.**

The `code-review-graph` project implements a **token-efficient** pipeline for generating review contexts. This system ensures that large language models receive only the relevant portions of a codebase when performing code reviews, dramatically reducing token consumption while preserving critical context. The entire flow is orchestrated through the public API `get_review_context_tool` defined in [`code_review_graph/main.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/main.py).

## Detect Changed Files

The review context generation flow begins with **change detection**. The `detect_changes_tool` function, implemented in [`code_review_graph/tools/analysis_tools.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/tools/analysis_tools.py), analyzes the current Git commit or pull request. It collects the complete set of files that were added, removed, or modified. Each file is annotated with exact line-range changes to establish precise boundaries for downstream processing.

This stage establishes the foundation for all subsequent context extraction. Without accurate change detection, the sub-graph construction would include irrelevant code paths or miss critical dependencies.

## Construct a Focused Sub-graph

With the list of changed files identified, the flow proceeds to **impact analysis**. The `get_review_context` function calls the graph-query engine in [`code_review_graph/graph.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/graph.py) to retrieve the minimal set of nodes and edges reachable from the changed files.

This **impact sub-graph** contains:

- Directly affected symbols
- Callers and callees of changed functions
- Cross-language dependencies (for example, Python ↔ JavaScript)

The sub-graph is then condensed into a hierarchical structure suitable for textual rendering. This condensation step is critical for token efficiency—it transforms raw graph data into a compact summary that retains structural relationships without verbose serialization.

## Gather Source Snippets and Assemble the Prompt

The final stage transforms the sub-graph into concrete review material. For each node in the sub-graph, `get_review_context` extracts relevant source excerpts surrounding the changed lines. These snippets are trimmed to fit within the LLM's context window constraints.

Each snippet receives metadata labeling:

- File path
- Line numbers
- Role description (for example, "function `process_data` calls `validate`")

The **review prompt** combines four components:

1. High-level change description
2. Condensed sub-graph summary
3. Curated source snippets
4. Fixed instruction block from [`code_review_graph/prompts.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/prompts.py)

The result is returned as a `ReviewContext` object ready for LLM consumption.

## Public Entry Point and Error Handling

The `get_review_context_tool` function wraps the core routine with provenance tracking and error handling. This design makes the flow safe to invoke from multiple interfaces:

- The CLI command `crg review`
- Programmatic usage via the MCP toolset
- Direct Python import

### Python API Usage

```python
from code_review_graph.main import get_review_context_tool

# Generate a review context for the current PR, using the default base branch.

review_ctx = get_review_context_tool(base="main")
print(review_ctx.prompt)        # Full LLM-ready prompt

print(review_ctx.snippets)      # Mapping of file → trimmed source excerpts

```

### CLI Usage

```bash

# Token-optimised review context

crg review --base main > review_prompt.txt

```

### Manual Integration

```python

# Manual integration in a custom script

changes = detect_changes_tool()
subgraph = build_subgraph(changes)
ctx = get_review_context(base="main", changes=changes, subgraph=subgraph)
llm_response = call_llm(ctx.prompt)

```

## Pipeline Architecture

The complete flow can be visualized as:

```

detect_changes_tool  →  sub-graph extraction  →  source-snippet collection
        │                     │                        │
        └───────► get_review_context_tool ◄───────────────────┘

```

This architecture enforces a strict separation of concerns while maintaining composability. Each stage consumes the output of the previous stage, enabling flexible integration into custom workflows.

## Key Source Files

The review context generation flow spans five essential files in the repository:

- [`code_review_graph/tools/review.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/tools/review.py) — Core implementation of `get_review_context`
- [`code_review_graph/main.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/main.py) — Public wrapper `get_review_context_tool`
- [`code_review_graph/graph.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/graph.py) — Graph data structures and sub-graph extraction logic
- [`code_review_graph/prompts.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/prompts.py) — Fixed instruction block prepended to generated prompts
- [`skills/review-pr/SKILL.md`](https://github.com/tirth8205/code-review-graph/blob/main/skills/review-pr/SKILL.md) — Example usage pattern for PR-review workflows

## Summary

- The review context generation flow uses three coordinated stages to minimize token usage while preserving review accuracy
- Change detection in [`analysis_tools.py`](https://github.com/tirth8205/code-review-graph/blob/main/analysis_tools.py) establishes precise boundaries for modified code
- Sub-graph extraction in [`graph.py`](https://github.com/tirth8205/code-review-graph/blob/main/graph.py) identifies only relevant dependent symbols across language boundaries
- Source snippet assembly applies windowing and metadata labeling to prepare LLM-ready content
- The `get_review_context_tool` wrapper provides safe, tracked access from CLI, API, or programmatic contexts

## Frequently Asked Questions

### What makes the review context generation flow token-efficient?

The flow achieves token efficiency through **selective inclusion**. Instead of passing entire files to the LLM, it constructs an impact sub-graph containing only symbols reachable from changed files. Source snippets are windowed around change points and condensed hierarchically. According to the `code-review-graph` source code, this typically reduces context size by 80-95% compared to full-file approaches.

### How does code-review-graph handle cross-language dependencies?

The graph engine in [`code_review_graph/graph.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/graph.py) maintains a unified symbol graph across multiple languages. When building the impact sub-graph, it follows edges regardless of source language—enabling detection of JavaScript callers when Python functions change. This is critical for modern polyglot codebases where review accuracy depends on understanding full call chains.

### Can I customize the prompt template used in review context generation?

The fixed instruction block lives in [`code_review_graph/prompts.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/prompts.py) and is currently prepended to all generated prompts. For template customization, you can modify this file directly or post-process the `ReviewContext.prompt` field after generation. The project structure separates prompt construction from core logic, enabling straightforward extension.

### What is the difference between `get_review_context` and `get_review_context_tool`?

`get_review_context` in [`code_review_graph/tools/review.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/tools/review.py) implements the core three-stage pipeline. `get_review_context_tool` in [`code_review_graph/main.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/main.py) adds a **provenance-tracking wrapper** with standardized error handling and MCP tool compatibility. For most use cases, import from [`main.py`](https://github.com/tirth8205/code-review-graph/blob/main/main.py); for custom orchestration requiring direct control, use the lower-level function from [`review.py`](https://github.com/tirth8205/code-review-graph/blob/main/review.py).