Code Review Graph Architecture: A Four-Layer Deep Dive into the Knowledge-First Codebase Engine

Code Review Graph uses a four-layer architecture—persistence, parsing, language resolution, and service APIs—to build and query a SQLite-backed knowledge graph of repository structure and relationships.

The Code Review Graph (CRG) is a modular, language-agnostic analysis engine that transforms source code into a queryable knowledge graph. According to the tirth8205/code-review-graph source, it trades heavy external databases for a lightweight SQLite core while maintaining the traversal performance needed for impact analysis and visualization at scale.

Architecture Overview: Four Layers Working in Concert

The codebase organizes functionality into four distinct layers that handle everything from raw file ingestion to semantic search and visualization.

Layer Purpose Core Files
Persistence SQLite-backed storage with batched transactions and NetworkX caching graph.py
Parsing & Normalisation File system walking, path normalization, and initial node/edge creation parser.py
Language Resolution Per-language symbol extraction and relationship mapping python_resolver.py, java_resolver.py, spring_resolver.py, cpp_scoped_resolver.py
Service & API CLI, daemon, search, impact analysis, flows, and visualization cli.py, daemon.py, search.py, flows.py, visualization.py, embeddings.py

Layer 1: The Persistence Layer (graph.py)

The GraphStore class in code_review_graph/graph.py serves as the single source of truth for all graph operations. It implements a durable, transactional storage system backed by SQLite.

Key design decisions here:

  • Atomic batches: All node and edge writes happen in transactions to prevent partial imports
  • NetworkX cache: An in-memory graph cache enables O(1) edge lookups during traversals, invalidated on writes
  • Schema versioning: The store manages its own migrations for forward compatibility

The core dataclasses GraphNode and GraphEdge define the graph's type system, capturing everything from files and functions to test relationships and inheritance chains.

from pathlib import Path
from code_review_graph.graph import GraphStore
from code_review_graph.parser import parse_file

db_path = Path(".crg/store.db")
with GraphStore(db_path) as store:
    file_path = "src/example.py"
    nodes, edges = parse_file(file_path)
    store.store_file_nodes_edges(file_path, nodes, edges, fhash="abc123")

Source: GraphStore.store_file_nodes_edges in graph.py (lines 59-71).

The persistence layer also handles post-import resolution through methods like resolve_bare_call_targets() and resolve_bare_tested_by_sources()—critical for cleaning up ambiguous symbol references after bulk imports.

Layer 2: Parsing & Normalisation (parser.py)

The Parser orchestrates the transformation from raw source files to structured graph data. Its responsibilities include:

  • Walking the file system according to include/exclude patterns
  • Selecting the appropriate language resolver based on file extensions
  • Normalizing paths and qualifying symbols
  • Creating "bare" edges for symbols that cannot be immediately resolved

The parser returns NodeInfo and EdgeInfo objects—lightweight containers that decouple extraction from storage. This separation allows resolvers to focus on language semantics while the parser handles universal concerns like path handling and batching.

Layer 3: Language Resolution

CRG ships with specialized resolvers for major language families, each understanding the grammar, scope rules, and module systems of its target.

Python Resolver (python_resolver.py)

Extracts functions, classes, and methods; resolves import statements to IMPORTS edges; maps call sites to CALLS edges. Handles Python's dynamic import patterns and relative imports.

Java/Spring Resolvers (java_resolver.py, spring_resolver.py)

The Java resolver handles standard Java classes and packages. The Spring resolver extends this with awareness of:

  • @Controller, @Service, @Repository, @Component annotations
  • Bean wiring relationships
  • HTTP endpoint detection from @RequestMapping and friends

This enables CRG to answer questions like "which API endpoints depend on this database repository?"

C++ and Additional Resolvers

The cpp_scoped_resolver.py (referenced as rescript_resolver.py in paths) and other language modules bring the same relationship extraction to native code, handling header includes, template instantiation, and namespace-scoped call resolution.

Cross-Language Compatibility

JS/TS/TSX files are treated as an interchangeable language family via _compatible_edge_languages, reducing false positives when calls cross syntax boundaries in the same codebase.

Layer 4: Service & API Layer

This layer exposes the graph's power through multiple interfaces.

Command-Line Interface (cli.py)

The crg entry point supports:

  • crg index — Full or incremental repository ingestion
  • crg search — Symbol lookup with FTS5-backed ranking
  • crg impact — Change impact radius analysis
  • crg viz — GraphViz DOT output generation

Incremental Daemon (daemon.py)

The daemon mode watches the repository for file changes and re-indexes modified files in background. This eliminates full re-parsing in CI environments and enables near-real-time graph updates during development.

Search & Impact Analysis (search.py, flows.py)

The search.py module wraps GraphStore.search_nodes() with higher-level APIs, while flows.py implements flow tracing and criticality scoring.

Impact radius queries—essential for code review automation—traverse the graph using edge-direction constants defined in constants.py:

from code_review_graph.graph import GraphStore
from code_review_graph.get_impact_radius import get_impact_radius

with GraphStore("mydb.db") as store:
    impacted = get_impact_radius(
        store,
        source_qn="src/services/user_service.py::deactivate_user",
        direction="outgoing",
        max_depth=3,
    )
    for qn in impacted:
        print(qn)

Source: get_impact_radius relies on edge-direction constants in constants.py.

Visualization (visualization.py)

The render_subgraph() function generates GraphViz DOT or interactive HTML for exploring code relationships:

from code_review_graph.graph import GraphStore
from code_review_graph.visualization import render_subgraph

with GraphStore("mydb.db") as store:
    dot = render_subgraph(store, root_qn="src/models/User.py::User")
    with open("user_graph.dot", "w") as f:
        f.write(dot)

Source: render_subgraph in visualization.py.

Semantic Search (embeddings.py)

Optional integration with vector embeddings enables semantic similarity search—finding code by meaning rather than exact symbol match. The system falls back to FTS5 when embeddings are unavailable, ensuring graceful degradation.

How the Layers Orchestrate: The Ingestion Pipeline

Understanding Code Review Graph architecture requires seeing how data flows through all four layers:

  1. CLI triggers indexing → cli.py invokes the parser
  2. Parser selects resolver → Based on file extension, hands file to appropriate language module
  3. Resolver extracts symbols → Returns NodeInfo/EdgeInfo with raw and qualified names
  4. GraphStore persists batch → SQLite transaction commits, NetworkX cache refreshes
  5. Post-processing resolves bare edges → Ambiguous references clarified using import evidence
  6. Queries traverse the graph → Impact radius, search, and visualization APIs consume the stored data

Key Design Decisions and Their Rationale

Decision Rationale
SQLite + NetworkX cache Durability without external database overhead; cache provides traversal speed
Evidence-backed bare-edge resolution Prevents accidental cross-repo name collisions by requiring import/file evidence
Two-phase search (FTS5 → LIKE) Fast tokenized search on modern schemas, graceful fallback on older databases
Language-family grouping Reduces false positives when calls cross JS/TS/TSX boundaries
Daemon incremental updates Maintains fresh graphs without full re-parsing, critical for CI efficiency

Summary

  • Code Review Graph architecture centers on a four-layer design: persistence, parsing, language resolution, and service APIs
  • The GraphStore in graph.py provides transactional SQLite storage with an in-memory NetworkX cache for performance
  • Language-specific resolvers in python_resolver.py, spring_resolver.py, and related files extract semantically meaningful relationships
  • The service layer exposes functionality through CLI, daemon, search, impact analysis, and visualization modules
  • Post-import resolution and evidence-backed edge linking ensure graph accuracy despite ambiguous source references

Frequently Asked Questions

What database does Code Review Graph use?

Code Review Graph uses SQLite as its primary persistence layer. This choice guarantees durability and fast random access without requiring external database infrastructure. The GraphStore class in graph.py manages all database operations, while a cached NetworkX graph provides O(1) edge lookups for in-memory traversals.

How does Code Review Graph handle multiple programming languages?

Each supported language has a dedicated resolver module that understands its grammar and module system. The parser.py orchestrator selects the appropriate resolver based on file extension—routing .py files to python_resolver.py, Spring Java files to spring_resolver.py, and so on. Resolvers return standardized NodeInfo and EdgeInfo objects that the graph stores uniformly.

What is "impact radius" in Code Review Graph?

Impact radius is a graph traversal feature that identifies all code potentially affected by a change. The get_impact_radius function (available through CLI as crg impact) walks outgoing or incoming edges from a starting symbol to a configurable depth, using edge-type and direction constants from constants.py. This enables automated risk assessment during code review.

How does Code Review Graph stay updated with code changes?

The daemon mode (daemon.py) runs as a background process that watches the repository filesystem, detects modified files, and re-indexes only those changes. This incremental approach keeps the graph current without the latency of full re-parsing, making it suitable for continuous integration pipelines and active development workflows.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →