Code Review Graph Architecture: A Four-Layer Deep Dive into the Knowledge-First Codebase Engine
Code Review Graph uses a four-layer architecture—persistence, parsing, language resolution, and service APIs—to build and query a SQLite-backed knowledge graph of repository structure and relationships.
The Code Review Graph (CRG) is a modular, language-agnostic analysis engine that transforms source code into a queryable knowledge graph. According to the tirth8205/code-review-graph source, it trades heavy external databases for a lightweight SQLite core while maintaining the traversal performance needed for impact analysis and visualization at scale.
Architecture Overview: Four Layers Working in Concert
The codebase organizes functionality into four distinct layers that handle everything from raw file ingestion to semantic search and visualization.
| Layer | Purpose | Core Files |
|---|---|---|
| Persistence | SQLite-backed storage with batched transactions and NetworkX caching | graph.py |
| Parsing & Normalisation | File system walking, path normalization, and initial node/edge creation | parser.py |
| Language Resolution | Per-language symbol extraction and relationship mapping | python_resolver.py, java_resolver.py, spring_resolver.py, cpp_scoped_resolver.py |
| Service & API | CLI, daemon, search, impact analysis, flows, and visualization | cli.py, daemon.py, search.py, flows.py, visualization.py, embeddings.py |
Layer 1: The Persistence Layer (graph.py)
The GraphStore class in code_review_graph/graph.py serves as the single source of truth for all graph operations. It implements a durable, transactional storage system backed by SQLite.
Key design decisions here:
- Atomic batches: All node and edge writes happen in transactions to prevent partial imports
- NetworkX cache: An in-memory graph cache enables O(1) edge lookups during traversals, invalidated on writes
- Schema versioning: The store manages its own migrations for forward compatibility
The core dataclasses GraphNode and GraphEdge define the graph's type system, capturing everything from files and functions to test relationships and inheritance chains.
from pathlib import Path
from code_review_graph.graph import GraphStore
from code_review_graph.parser import parse_file
db_path = Path(".crg/store.db")
with GraphStore(db_path) as store:
file_path = "src/example.py"
nodes, edges = parse_file(file_path)
store.store_file_nodes_edges(file_path, nodes, edges, fhash="abc123")
Source: GraphStore.store_file_nodes_edges in graph.py (lines 59-71).
The persistence layer also handles post-import resolution through methods like resolve_bare_call_targets() and resolve_bare_tested_by_sources()—critical for cleaning up ambiguous symbol references after bulk imports.
Layer 2: Parsing & Normalisation (parser.py)
The Parser orchestrates the transformation from raw source files to structured graph data. Its responsibilities include:
- Walking the file system according to include/exclude patterns
- Selecting the appropriate language resolver based on file extensions
- Normalizing paths and qualifying symbols
- Creating "bare" edges for symbols that cannot be immediately resolved
The parser returns NodeInfo and EdgeInfo objects—lightweight containers that decouple extraction from storage. This separation allows resolvers to focus on language semantics while the parser handles universal concerns like path handling and batching.
Layer 3: Language Resolution
CRG ships with specialized resolvers for major language families, each understanding the grammar, scope rules, and module systems of its target.
Python Resolver (python_resolver.py)
Extracts functions, classes, and methods; resolves import statements to IMPORTS edges; maps call sites to CALLS edges. Handles Python's dynamic import patterns and relative imports.
Java/Spring Resolvers (java_resolver.py, spring_resolver.py)
The Java resolver handles standard Java classes and packages. The Spring resolver extends this with awareness of:
@Controller,@Service,@Repository,@Componentannotations- Bean wiring relationships
- HTTP endpoint detection from
@RequestMappingand friends
This enables CRG to answer questions like "which API endpoints depend on this database repository?"
C++ and Additional Resolvers
The cpp_scoped_resolver.py (referenced as rescript_resolver.py in paths) and other language modules bring the same relationship extraction to native code, handling header includes, template instantiation, and namespace-scoped call resolution.
Cross-Language Compatibility
JS/TS/TSX files are treated as an interchangeable language family via _compatible_edge_languages, reducing false positives when calls cross syntax boundaries in the same codebase.
Layer 4: Service & API Layer
This layer exposes the graph's power through multiple interfaces.
Command-Line Interface (cli.py)
The crg entry point supports:
crg index— Full or incremental repository ingestioncrg search— Symbol lookup with FTS5-backed rankingcrg impact— Change impact radius analysiscrg viz— GraphViz DOT output generation
Incremental Daemon (daemon.py)
The daemon mode watches the repository for file changes and re-indexes modified files in background. This eliminates full re-parsing in CI environments and enables near-real-time graph updates during development.
Search & Impact Analysis (search.py, flows.py)
The search.py module wraps GraphStore.search_nodes() with higher-level APIs, while flows.py implements flow tracing and criticality scoring.
Impact radius queries—essential for code review automation—traverse the graph using edge-direction constants defined in constants.py:
from code_review_graph.graph import GraphStore
from code_review_graph.get_impact_radius import get_impact_radius
with GraphStore("mydb.db") as store:
impacted = get_impact_radius(
store,
source_qn="src/services/user_service.py::deactivate_user",
direction="outgoing",
max_depth=3,
)
for qn in impacted:
print(qn)
Source: get_impact_radius relies on edge-direction constants in constants.py.
Visualization (visualization.py)
The render_subgraph() function generates GraphViz DOT or interactive HTML for exploring code relationships:
from code_review_graph.graph import GraphStore
from code_review_graph.visualization import render_subgraph
with GraphStore("mydb.db") as store:
dot = render_subgraph(store, root_qn="src/models/User.py::User")
with open("user_graph.dot", "w") as f:
f.write(dot)
Source: render_subgraph in visualization.py.
Semantic Search (embeddings.py)
Optional integration with vector embeddings enables semantic similarity search—finding code by meaning rather than exact symbol match. The system falls back to FTS5 when embeddings are unavailable, ensuring graceful degradation.
How the Layers Orchestrate: The Ingestion Pipeline
Understanding Code Review Graph architecture requires seeing how data flows through all four layers:
- CLI triggers indexing →
cli.pyinvokes the parser - Parser selects resolver → Based on file extension, hands file to appropriate language module
- Resolver extracts symbols → Returns
NodeInfo/EdgeInfowith raw and qualified names - GraphStore persists batch → SQLite transaction commits, NetworkX cache refreshes
- Post-processing resolves bare edges → Ambiguous references clarified using import evidence
- Queries traverse the graph → Impact radius, search, and visualization APIs consume the stored data
Key Design Decisions and Their Rationale
| Decision | Rationale |
|---|---|
| SQLite + NetworkX cache | Durability without external database overhead; cache provides traversal speed |
| Evidence-backed bare-edge resolution | Prevents accidental cross-repo name collisions by requiring import/file evidence |
| Two-phase search (FTS5 → LIKE) | Fast tokenized search on modern schemas, graceful fallback on older databases |
| Language-family grouping | Reduces false positives when calls cross JS/TS/TSX boundaries |
| Daemon incremental updates | Maintains fresh graphs without full re-parsing, critical for CI efficiency |
Summary
- Code Review Graph architecture centers on a four-layer design: persistence, parsing, language resolution, and service APIs
- The
GraphStoreingraph.pyprovides transactional SQLite storage with an in-memory NetworkX cache for performance - Language-specific resolvers in
python_resolver.py,spring_resolver.py, and related files extract semantically meaningful relationships - The service layer exposes functionality through CLI, daemon, search, impact analysis, and visualization modules
- Post-import resolution and evidence-backed edge linking ensure graph accuracy despite ambiguous source references
Frequently Asked Questions
What database does Code Review Graph use?
Code Review Graph uses SQLite as its primary persistence layer. This choice guarantees durability and fast random access without requiring external database infrastructure. The GraphStore class in graph.py manages all database operations, while a cached NetworkX graph provides O(1) edge lookups for in-memory traversals.
How does Code Review Graph handle multiple programming languages?
Each supported language has a dedicated resolver module that understands its grammar and module system. The parser.py orchestrator selects the appropriate resolver based on file extension—routing .py files to python_resolver.py, Spring Java files to spring_resolver.py, and so on. Resolvers return standardized NodeInfo and EdgeInfo objects that the graph stores uniformly.
What is "impact radius" in Code Review Graph?
Impact radius is a graph traversal feature that identifies all code potentially affected by a change. The get_impact_radius function (available through CLI as crg impact) walks outgoing or incoming edges from a starting symbol to a configurable depth, using edge-type and direction constants from constants.py. This enables automated risk assessment during code review.
How does Code Review Graph stay updated with code changes?
The daemon mode (daemon.py) runs as a background process that watches the repository filesystem, detects modified files, and re-indexes only those changes. This incremental approach keeps the graph current without the latency of full re-parsing, making it suitable for continuous integration pipelines and active development workflows.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →