How Code Review Graph (CRG) Reduces Token Usage for AI Coding Assistants: A 65× Efficiency Breakdown

Code Review Graph (CRG) reduces token usage for AI coding assistants by converting codebases into structured knowledge graphs and serving only the minimal context slice required for each review task—achieving median token reductions of 65× and up to 376× in benchmarked repositories.

AI coding assistants face a fundamental constraint: context windows are expensive and finite. Sending an entire codebase to a large language model (LLM) quickly exhausts token budgets, drives up API costs, and degrades response quality. CRG, an open-source tool in the tirth8205/code-review-graph repository, solves this by replacing bulk source code transmission with precise, graph-queried context retrieval.

Parse Repositories into Compact AST Representations

The first optimization layer happens at ingestion. Instead of preserving raw source text, CRG uses Tree-sitter to parse every file into an abstract syntax tree (AST).

In code_review_graph/parser.py, the parser extracts entities—functions, classes, imports, and call relationships—into a language-agnostic format. This AST representation is dramatically more compact than source code because it strips comments, formatting, and redundant syntax while preserving semantic structure.

The result: a dense knowledge graph where nodes represent code entities and edges represent relationships like "calls," "imports," or "contains."

Store and Query via SQLite for Indexed Retrieval

Once parsed, the graph persists in SQLite through code_review_graph/backend/sqlite.py (orchestrated via code_review_graph/graph.py).

This storage strategy enables fast, indexed lookups that would be impossible with file-system scanning. When the AI needs context, CRG queries the graph database rather than reading and tokenizing raw files. The AI never sees the full repository—only the specific nodes and edges retrieved by the query.

Compute Precise Blast-Radius Impact Sets

The core token-saving mechanism is incremental blast-radius analysis. When a file changes, CRG does not resubmit the entire repository or even the full diff.

Instead, the graph engine walks caller/callee and import edges to compute the impact set—the transitive closure of code affected by the change. This includes:

  • Changed files themselves
  • Direct and indirect callers
  • Dependent tests
  • Related configuration

As illustrated in the README's blast-radius analysis section, this graph traversal isolates exactly what needs review while excluding unrelated code. The blast-radius is typically under 5% of the total codebase, even for substantial changes.

MCP Tools Return Minimal Context with Savings Metadata

CRG exposes its functionality through Model Context Protocol (MCP) tools designed for token-conscious AI interactions:

  • get_minimal_context_tool – Returns the smallest viable context for a query
  • get_review_context_tool – Fetches review-relevant snippets with risk scoring
  • get_impact_radius_tool – Returns blast-radius nodes for a given change

Each tool response in code_review_graph/context_savings.py includes a context_savings estimate: original token count, graph-derived token count, and percentage saved. Typical responses contain a few hundred tokens versus thousands for whole-corpus reads.

Verify Savings with Built-In Token Accounting

CRG makes token reduction transparent and verifiable:

  • Token Savings panel: CLI commands like detect-changes --brief and update --brief display a formatted panel showing:

    • Full context token count (e.g., 12,921 tokens)
    • Graph context used (e.g., 762 tokens)
    • Saved tokens with percentage (e.g., ~94%)
  • Optional verification: The --verify flag cross-checks estimates against OpenAI's cl100k_base tokenizer. Calibration data in docs/REPRODUCING.md confirms estimates stay within ~1% of actual GPT-4 token counts.

Benchmark Results: 65× Median Reduction

Empirical validation comes from code_review_graph/token_benchmark.py. Across six real repositories, CRG achieves:

  • Median token reduction: ~65×
  • Maximum reduction: 376× (largest repository)

These ratios represent real-world questions answered with graph-derived context versus full-corpus submission.

Practical Usage Flow


# Build the graph once (≈10s for 500 files)

code-review-graph build

# Or enable live updates

code-review-graph watch &

# Review latest change with token savings displayed

code-review-graph detect-changes --brief

Example output:


┌─────────────────────── Token Savings ────────────────────────┐
│ Full context would be:     12,921 tokens                     │
│ Graph context used:           762 tokens                     │
│ Saved:                     12,159 tokens (~94%)              │
└──────────────────────────────────────────────────────────────┘

Verify accuracy:

code-review-graph detect-changes --brief --verify

Python MCP client usage:

from mcp import Client

client = Client(command=["code-review-graph", "serve"])
response = client.run_tool(
    "get_review_context_tool",
    {"query": "How does authentication work?", "max_tokens": 2000}
)
print(response["context_savings"])

# {'original_tokens': 12000, 'saved_tokens': 11500, 'reduction_ratio': 6.3}

Key Architectural Components

File Purpose
code_review_graph/parser.py Tree-sitter AST extraction
code_review_graph/graph.py Graph construction and SQLite persistence
code_review_graph/context_savings.py Token estimation and savings formatting
code_review_graph/tools/review.py MCP tool implementations
code_review_graph/cli.py CLI with Token Savings panel
code_review_graph/token_benchmark.py Efficiency benchmarking suite

Summary

  • AST-level granularity replaces verbose source code with compact semantic entities
  • Edge-based blast-radius analysis limits context to genuinely impacted code
  • SQLite-backed graph queries enable millisecond retrieval without file I/O
  • MCP-native tools let AI assistants request exactly what they need
  • Verified token accounting proves 65× median reductions with ~1% estimate accuracy
  • Multi-hop query capability answers complex questions without repeated corpus scans

Frequently Asked Questions

How does CRG compare to simple file-grep for context reduction?

File-grep returns text matches without understanding semantic relationships. CRG's graph traversals follow actual code dependencies—calls, inheritance, imports—ensuring the AI receives logically complete context rather than syntactically similar but semantically irrelevant files. This precision is why CRG achieves 65× reduction where grep-based filtering often yields 2–5× at best.

Can CRG's token estimates be trusted for billing calculations?

Yes. The --verify flag in code_review_graph/context_savings.py validates estimates against OpenAI's official cl100k_base tokenizer. Documentation in docs/REPRODUCING.md shows calibrated errors under 1% across diverse codebases, making the estimates reliable for cost projection.

What programming languages does CRG support?

Tree-sitter parsers in code_review_graph/parser.py provide broad language coverage including Python, JavaScript, TypeScript, Go, Rust, Java, and C/C++. New languages are supported by adding their Tree-sitter grammar without changing the graph construction logic.

Does CRG work with self-hosted or non-OpenAI models?

Absolutely. CRG's MCP interface is model-agnostic. The token estimation uses cl100k_base for verification, but the graph-served context works with any LLM—Claude, Gemini, local models via Ollama, or enterprise deployments. The savings ratios remain valid as they're derived from context size reduction, not model-specific behavior.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →