What Are the Limitations of the Current Code-Review-Graph Implementation?

The code-review-graph implementation faces five critical constraints: circular benchmark metrics that inflate recall scores, structural overhead that makes trivial edits less efficient than naive file reads, low semantic search quality (MRR ≈ 0.35), flow detection limited to ~33% recall with weak JavaScript and Go support, and inherent precision-recall trade-offs that generate false positives in large dependency graphs.

Code Review Graph (CRG) is a language-agnostic, AST-based knowledge graph that parses repositories using Tree-sitter and exposes analysis via the Model Context Protocol (MCP). While it delivers dramatic token savings for complex changes, the current implementation in tirth8205/code-review-graph has documented architectural limitations that affect benchmarking accuracy, performance on small edits, and analysis precision across different programming languages.

Circular Benchmark Metrics in Impact Analysis

The most significant methodological limitation concerns how CRG measures the accuracy of its impact analysis. The system reports graph-derived recall = 1.0, but this metric is circular because the ground truth used for validation is produced from the same graph edges that the predictor traverses.

In code_review_graph/tools/analysis_tools.py, the blast-radius algorithm walks caller, callee, and test edges to identify potentially impacted nodes. However, because the validation set derives from these identical edges, the recall measurement represents an upper bound rather than a true independent verification of predictive accuracy. This limitation is explicitly documented in the README's Limitations section (lines 81-88), warning that the recall metric should be interpreted as a theoretical maximum rather than empirical performance.

Performance Overhead on Trivial Changes

CRG incurs structural metadata costs that make it less efficient than naive file reads for trivial edits. The system uses incremental updates—re-parsing only files whose SHA-256 hashes changed—but even these updates require loading graph metadata from the SQLite database.

When you run a small single-file change through the pipeline defined in code_review_graph/parser.py, the overhead of node-type mapping, edge generation, and database lookups can produce a negative token savings count. As noted in the CLI output from code_review_graph/main.py, the "Saved" token metric may show that the graph context exceeds the raw file size for isolated modifications.


# Detect changes showing negative token savings for trivial edits

code-review-graph detect-changes --brief

# → The "Saved" token count may be negative because structural metadata 

#   outweighs the raw file size for small changes.

Suboptimal Search Quality and Ranking

The semantic search capability in code_review_graph/tools/query.py achieves a Mean Reciprocal Rank (MRR) of approximately 0.35, indicating that the correct result often appears outside the top-ranked position. While keyword search frequently returns the correct file within the top-4 results, the ranking algorithm requires improvement.

This limitation manifests differently across project types. Some JavaScript projects like Express experience zero hits for certain queries, demonstrating that the current heuristics in the query implementation struggle with specific codebases despite the language-agnostic Tree-sitter foundation.

Limited Flow Detection Coverage

Flow detection—the identification of data flow through function calls and variable assignments—operates at approximately 33% recall in the current implementation. The code_review_graph/tools/analysis_tools.py file contains heuristics that work best for Python and PHP/Laravel codebases, while JavaScript and Go support remains underdeveloped.

This language disparity means that control flow analysis and taint tracking are significantly more reliable in Python repositories than in Node.js or Go projects, limiting the utility of security-focused reviews in polyglot environments.

Precision-Recall Trade-offs in Blast Radius

The impact analysis system is intentionally conservative, configured via environment variables CRG_MAX_IMPACT_NODES and CRG_MAX_IMPACT_DEPTH to prevent unbounded graph traversals. However, this conservatism creates a precision-recall trade-off where the algorithm flags might-be-affected files to avoid missing true dependencies.

In large dependency graphs, this approach generates false positives, returning extensive lists of potentially impacted nodes when only a subset actually requires review. You can observe this behavior when running the impact radius tool on highly connected files:


# Examine blast radius showing conservative flagging

code-review-graph get-impact-radius --file src/app/login.py

# → Returns many nodes including false positives due to conservative 

#   traversal limits defined in CRG_MAX_IMPACT_NODES.

Mitigation Strategies

You can partially address these limitations through configuration. The CRG_TOOLS environment variable allows you to trim the default set of ~30 MCP tools to reduce token usage and complexity, though this requires explicitly enabling specific capabilities like flow detection or impact analysis.


# Limit toolset to reduce overhead in token-constrained environments

CRG_TOOLS=query_graph_tool,detect_changes_tool \
code-review-graph serve

# → Only selected tools load, but you lose flow-detection and 

#   impact-radius capabilities unless explicitly enabled.

For the circular benchmark issue, treat the recall = 1.0 metric as a sanity check rather than a performance target. When working with small changes, consider bypassing the graph entirely for single-file edits under 50 lines to avoid the metadata overhead documented in code_review_graph/parser.py.

Summary

  • Circular metrics: The recall = 1.0 score is graph-derived and therefore represents an upper bound rather than independent validation.
  • Overhead on small edits: Incremental updates in parser.py create metadata costs that exceed naive file reads for trivial changes.
  • Low search MRR: Semantic search achieves only ~0.35 MRR, with some projects experiencing zero relevant hits.
  • Flow detection gaps: 33% recall concentrated in Python/PHP, with JavaScript and Go requiring improved heuristics in analysis_tools.py.
  • False positive rates: Conservative blast-radius analysis in query.py generates false positives in large dependency graphs due to the precision-recall trade-off.

Frequently Asked Questions

Why is the recall metric considered circular?

The recall metric is circular because the ground truth used to validate impact analysis is generated from the same graph edges that the prediction algorithm traverses. According to the README documentation (lines 81-88), this means the "recall = 1.0" measurement validates that the graph can find its own edges rather than proving predictive accuracy against external validation data.

When is code-review-graph less efficient than reading files directly?

CRG becomes less efficient for single-file changes involving fewer than approximately 50 lines of code. In these cases, the overhead of loading structural metadata from the SQLite database and parsing AST nodes in code_review_graph/parser.py produces a larger context window than simply reading the raw file contents.

Which programming languages have the best flow detection support?

Python and PHP/Laravel receive the most reliable flow detection support, while JavaScript and Go exhibit significantly weaker heuristics. The current implementation achieves approximately 33% recall overall, with the code_review_graph/tools/analysis_tools.py file optimized primarily for Pythonic control flow patterns.

How can I reduce token usage when using the MCP server?

Set the CRG_TOOLS environment variable to limit the default set of ~30 MCP tools to only those required for your specific task. For example, restricting the toolset to query_graph_tool and detect_changes_tool reduces initialization overhead, though you must explicitly add get_impact_radius_tool or flow detection tools if you need blast-radius analysis.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →