# Limitations of Code Review Graph: 5 Critical Constraints Explained

> Discover the 5 critical limitations of Code Review Graph including circular impact metrics, low search accuracy, and limited flow detection. Understand its constraints for better analysis.

- Repository: [Tirth Kanani/code-review-graph](https://github.com/tirth8205/code-review-graph)
- Tags: deep-dive
- Published: 2026-08-10

---

**Code Review Graph suffers from circular impact metrics, inefficiency on trivial edits, low search ranking accuracy (MRR ≈ 0.35), limited flow detection recall (~33%), and conservative precision-recall trade-offs that generate false positives in dense dependency graphs.**

Code Review Graph (CRG) is an open-source tool that constructs a language-agnostic knowledge graph from Tree-sitter ASTs and exposes repository structure via the Model Context Protocol (MCP). While it delivers dramatic token savings for complex multi-file changes, understanding the limitations of Code Review Graph is essential for teams evaluating its suitability for high-precision workflows or small-scale edits.

## Circular Ground-Truth in Impact Analysis

The most significant methodological limitation concerns the **graph-derived recall metric**. CRG claims a "recall = 1.0" for impact analysis, but this measurement is circular because the ground-truth is produced from the same graph edges the predictor traverses.

According to the [README.md](https://github.com/tirth8205/code-review-graph/blob/main/README.md#L81-L88) limitations section, this creates an upper-bound measurement rather than a true recall validation against external test suites. When [`code_review_graph/tools/analysis_tools.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/tools/analysis_tools.py) computes the blast radius, it relies on caller, callee, and test edges stored in the local SQLite database. Consequently, the system cannot detect impacts from relationships it failed to parse during the initial AST-to-graph conversion in [`code_review_graph/parser.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/parser.py).

## Overhead on Trivial Changes

For small, single-file modifications, CRG can be **less efficient than naive file reads**. The overhead of loading structural metadata from the graph database often exceeds the token cost of simply reading the raw file content.

This limitation surfaces when running:

```bash
code-review-graph detect-changes --brief

```

For trivial edits, the "Saved" token count may display negative values because the graph context includes AST node types, edge relationships, and SHA-256 hash metadata that outweigh the actual source code bytes. The incremental update system—implemented in [`code_review_graph/main.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/main.py)—re-parses only changed files via hash comparison, but the fixed cost of graph metadata retrieval remains constant regardless of edit size.

## Search Quality and Ranking Limitations

CRG's keyword search exhibits a **Mean Reciprocal Rank (MRR) of approximately 0.35**, indicating that the correct result often appears within the top-4 hits but ranking accuracy degrades quickly beyond that threshold.

The search implementation in [`code_review_graph/tools/query.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/tools/query.py) struggles with certain project structures; for example, the Express.js repository reportedly returns zero hits for some queries. When executing:

```bash
code-review-graph query "how to authenticate"

```

Results may require scanning multiple entries due to the current ranking heuristics, which lack semantic understanding beyond graph proximity metrics.

## Flow Detection Coverage Gaps

The **flow detection recall stands at roughly 33%**, with significant language disparities. The current heuristics work reliably for Python and PHP/Laravel codebases but fail to capture control flow adequately in JavaScript and Go projects.

This limitation stems from the parser configuration in [`code_review_graph/parser.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/parser.py), where Tree-sitter node type mappings differ across languages. Complex asynchronous patterns in JavaScript or goroutine concurrency in Go generate AST structures that the current flow analysis algorithms in [`analysis_tools.py`](https://github.com/tirth8205/code-review-graph/blob/main/analysis_tools.py) do not fully traverse.

## Precision vs. Recall Trade-offs

CRG's impact analysis is **intentionally conservative**, flagging any file that *might* be affected by a change. This approach maximizes recall (within the graph's constraints) at the expense of precision, generating false positives in repositories with dense dependency graphs.

When invoking:

```bash
code-review-graph get-impact-radius --file src/app/login.py

```

The tool may return an inflated set of impacted nodes due to the `CRG_MAX_IMPACT_NODES` and `CRG_MAX_IMPACT_DEPTH` environment variable limits. These constraints—documented in the README—prevent infinite traversal but can include tangentially related files that do not actually require review.

## Architectural Root Causes

Understanding why these limitations persist requires examining the core pipeline implementation.

### AST-to-Graph Pipeline Constraints

The conversion from Tree-sitter AST to graph nodes happens in [`code_review_graph/parser.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/parser.py), which maps functions, classes, imports, and call sites into a SQLite database. This process loses source-code ordering information and macro-expanded context, contributing to the flow detection gaps in compiled languages.

### Blast-Radius Algorithm Limits

The impact calculation relies on environment variables `CRG_MAX_IMPACT_NODES` and `CRG_MAX_IMPACT_DEPTH` to prevent runaway traversals. As implemented in [`code_review_graph/tools/query.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/tools/query.py), these hard limits can truncate legitimate dependency chains while still including false positives from shared utility nodes.

### MCP Tool Overhead

CRG exposes approximately 30 MCP tools by default, including `get_impact_radius_tool` and `detect_changes_tool`. While you can restrict the toolset via the `CRG_TOOLS` environment variable:

```bash
CRG_TOOLS=query_graph_tool,detect_changes_tool code-review-graph serve

```

This reduces token usage but eliminates flow-detection and impact-radius capabilities unless explicitly re-enabled, creating a configuration burden for resource-constrained environments.

## Summary

- **Circular metrics**: The recall = 1.0 claim is graph-derived and represents an upper bound, not true external validation.
- **Trivial edit overhead**: Small changes incur fixed graph metadata costs that can exceed raw file reads.
- **Search ranking**: MRR ≈ 0.35 indicates keyword search requires top-4 scanning for reliable results.
- **Flow detection**: 33% recall with poor JavaScript and Go support due to AST traversal limitations.
- **False positives**: Conservative impact analysis favors recall over precision, inflating blast-radius results in dense codebases.

## Frequently Asked Questions

### Why does Code Review Graph show negative token savings for small edits?

The graph stores structural metadata—AST node types, edge relationships, and SHA-256 hashes—in a SQLite database. For trivial edits, retrieving this metadata costs more tokens than reading the raw file directly, causing negative savings in the `detect-changes` output.

### How accurate is the impact analysis in Code Review Graph?

Impact analysis achieves high recall within the graph's known edges but suffers from false positives due to conservative traversal. The system flags all potentially affected files, including those with tenuous relationships, to ensure no impacts are missed. The claimed "recall = 1.0" is circular because it validates against the same graph edges used for prediction.

### Which programming languages work best with Code Review Graph?

Python and PHP/Laravel exhibit the highest flow detection accuracy. JavaScript and Go show significantly lower recall (~33% overall) due to insufficient heuristics for asynchronous patterns and concurrency primitives in the current Tree-sitter parsing implementation.

### Can I reduce the tool overhead when running the MCP server?

Yes. Set the `CRG_TOOLS` environment variable to a comma-separated list of required tools (e.g., `query_graph_tool,detect_changes_tool`) before running `code-review-graph serve`. This reduces token consumption but disables excluded capabilities like flow detection unless explicitly included in the list.