# Code Review Graph Architecture: A Four-Layer Deep Dive into the Knowledge-First Codebase Engine

> Explore the four-layer architecture of Code Review Graph. Understand its persistence, parsing, language resolution, and API layers for querying repository knowledge.

- Repository: [Tirth Kanani/code-review-graph](https://github.com/tirth8205/code-review-graph)
- Tags: architecture
- Published: 2026-08-10

---

**Code Review Graph uses a four-layer architecture—persistence, parsing, language resolution, and service APIs—to build and query a SQLite-backed knowledge graph of repository structure and relationships.**

The **Code Review Graph** (CRG) is a modular, language-agnostic analysis engine that transforms source code into a queryable knowledge graph. According to the `tirth8205/code-review-graph` source, it trades heavy external databases for a lightweight SQLite core while maintaining the traversal performance needed for impact analysis and visualization at scale.

## Architecture Overview: Four Layers Working in Concert

The codebase organizes functionality into four distinct layers that handle everything from raw file ingestion to semantic search and visualization.

| Layer | Purpose | Core Files |
|-------|---------|-----------|
| **Persistence** | SQLite-backed storage with batched transactions and NetworkX caching | [`graph.py`](https://github.com/tirth8205/code-review-graph/blob/main/graph.py) |
| **Parsing & Normalisation** | File system walking, path normalization, and initial node/edge creation | [`parser.py`](https://github.com/tirth8205/code-review-graph/blob/main/parser.py) |
| **Language Resolution** | Per-language symbol extraction and relationship mapping | [`python_resolver.py`](https://github.com/tirth8205/code-review-graph/blob/main/python_resolver.py), [`java_resolver.py`](https://github.com/tirth8205/code-review-graph/blob/main/java_resolver.py), [`spring_resolver.py`](https://github.com/tirth8205/code-review-graph/blob/main/spring_resolver.py), [`cpp_scoped_resolver.py`](https://github.com/tirth8205/code-review-graph/blob/main/cpp_scoped_resolver.py) |
| **Service & API** | CLI, daemon, search, impact analysis, flows, and visualization | [`cli.py`](https://github.com/tirth8205/code-review-graph/blob/main/cli.py), [`daemon.py`](https://github.com/tirth8205/code-review-graph/blob/main/daemon.py), [`search.py`](https://github.com/tirth8205/code-review-graph/blob/main/search.py), [`flows.py`](https://github.com/tirth8205/code-review-graph/blob/main/flows.py), [`visualization.py`](https://github.com/tirth8205/code-review-graph/blob/main/visualization.py), [`embeddings.py`](https://github.com/tirth8205/code-review-graph/blob/main/embeddings.py) |

## Layer 1: The Persistence Layer (graph.py)

The **GraphStore** class in [`code_review_graph/graph.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/graph.py) serves as the single source of truth for all graph operations. It implements a durable, transactional storage system backed by SQLite.

Key design decisions here:

- **Atomic batches**: All node and edge writes happen in transactions to prevent partial imports
- **NetworkX cache**: An in-memory graph cache enables O(1) edge lookups during traversals, invalidated on writes
- **Schema versioning**: The store manages its own migrations for forward compatibility

The core dataclasses `GraphNode` and `GraphEdge` define the graph's type system, capturing everything from files and functions to test relationships and inheritance chains.

```python
from pathlib import Path
from code_review_graph.graph import GraphStore
from code_review_graph.parser import parse_file

db_path = Path(".crg/store.db")
with GraphStore(db_path) as store:
    file_path = "src/example.py"
    nodes, edges = parse_file(file_path)
    store.store_file_nodes_edges(file_path, nodes, edges, fhash="abc123")

```

*Source*: `GraphStore.store_file_nodes_edges` in [`graph.py`](https://github.com/tirth8205/code-review-graph/blob/main/graph.py) (lines 59-71).

The persistence layer also handles **post-import resolution** through methods like `resolve_bare_call_targets()` and `resolve_bare_tested_by_sources()`—critical for cleaning up ambiguous symbol references after bulk imports.

## Layer 2: Parsing & Normalisation (parser.py)

The **Parser** orchestrates the transformation from raw source files to structured graph data. Its responsibilities include:

- Walking the file system according to include/exclude patterns
- Selecting the appropriate **language resolver** based on file extensions
- Normalizing paths and qualifying symbols
- Creating "bare" edges for symbols that cannot be immediately resolved

The parser returns `NodeInfo` and `EdgeInfo` objects—lightweight containers that decouple extraction from storage. This separation allows resolvers to focus on language semantics while the parser handles universal concerns like path handling and batching.

## Layer 3: Language Resolution

CRG ships with **specialized resolvers** for major language families, each understanding the grammar, scope rules, and module systems of its target.

### Python Resolver (python_resolver.py)

Extracts functions, classes, and methods; resolves `import` statements to `IMPORTS` edges; maps call sites to `CALLS` edges. Handles Python's dynamic import patterns and relative imports.

### Java/Spring Resolvers (java_resolver.py, spring_resolver.py)

The Java resolver handles standard Java classes and packages. The **Spring resolver** extends this with awareness of:

- `@Controller`, `@Service`, `@Repository`, `@Component` annotations
- Bean wiring relationships
- HTTP endpoint detection from `@RequestMapping` and friends

This enables CRG to answer questions like "which API endpoints depend on this database repository?"

### C++ and Additional Resolvers

The [`cpp_scoped_resolver.py`](https://github.com/tirth8205/code-review-graph/blob/main/cpp_scoped_resolver.py) (referenced as [`rescript_resolver.py`](https://github.com/tirth8205/code-review-graph/blob/main/rescript_resolver.py) in paths) and other language modules bring the same relationship extraction to native code, handling header includes, template instantiation, and namespace-scoped call resolution.

### Cross-Language Compatibility

JS/TS/TSX files are treated as an interchangeable language family via `_compatible_edge_languages`, reducing false positives when calls cross syntax boundaries in the same codebase.

## Layer 4: Service & API Layer

This layer exposes the graph's power through multiple interfaces.

### Command-Line Interface (cli.py)

The `crg` entry point supports:

- `crg index` — Full or incremental repository ingestion
- `crg search` — Symbol lookup with FTS5-backed ranking
- `crg impact` — Change impact radius analysis
- `crg viz` — GraphViz DOT output generation

### Incremental Daemon (daemon.py)

The **daemon mode** watches the repository for file changes and re-indexes modified files in background. This eliminates full re-parsing in CI environments and enables near-real-time graph updates during development.

### Search & Impact Analysis (search.py, flows.py)

The [`search.py`](https://github.com/tirth8205/code-review-graph/blob/main/search.py) module wraps `GraphStore.search_nodes()` with higher-level APIs, while [`flows.py`](https://github.com/tirth8205/code-review-graph/blob/main/flows.py) implements flow tracing and criticality scoring.

Impact radius queries—essential for code review automation—traverse the graph using edge-direction constants defined in [`constants.py`](https://github.com/tirth8205/code-review-graph/blob/main/constants.py):

```python
from code_review_graph.graph import GraphStore
from code_review_graph.get_impact_radius import get_impact_radius

with GraphStore("mydb.db") as store:
    impacted = get_impact_radius(
        store,
        source_qn="src/services/user_service.py::deactivate_user",
        direction="outgoing",
        max_depth=3,
    )
    for qn in impacted:
        print(qn)

```

*Source*: `get_impact_radius` relies on edge-direction constants in [`constants.py`](https://github.com/tirth8205/code-review-graph/blob/main/constants.py).

### Visualization (visualization.py)

The `render_subgraph()` function generates GraphViz DOT or interactive HTML for exploring code relationships:

```python
from code_review_graph.graph import GraphStore
from code_review_graph.visualization import render_subgraph

with GraphStore("mydb.db") as store:
    dot = render_subgraph(store, root_qn="src/models/User.py::User")
    with open("user_graph.dot", "w") as f:
        f.write(dot)

```

*Source*: `render_subgraph` in [`visualization.py`](https://github.com/tirth8205/code-review-graph/blob/main/visualization.py).

### Semantic Search (embeddings.py)

Optional integration with vector embeddings enables **semantic similarity search**—finding code by meaning rather than exact symbol match. The system falls back to FTS5 when embeddings are unavailable, ensuring graceful degradation.

## How the Layers Orchestrate: The Ingestion Pipeline

Understanding **Code Review Graph architecture** requires seeing how data flows through all four layers:

1. **CLI triggers indexing** → [`cli.py`](https://github.com/tirth8205/code-review-graph/blob/main/cli.py) invokes the parser
2. **Parser selects resolver** → Based on file extension, hands file to appropriate language module
3. **Resolver extracts symbols** → Returns `NodeInfo`/`EdgeInfo` with raw and qualified names
4. **GraphStore persists batch** → SQLite transaction commits, NetworkX cache refreshes
5. **Post-processing resolves bare edges** → Ambiguous references clarified using import evidence
6. **Queries traverse the graph** → Impact radius, search, and visualization APIs consume the stored data

## Key Design Decisions and Their Rationale

| Decision | Rationale |
|----------|-----------|
| **SQLite + NetworkX cache** | Durability without external database overhead; cache provides traversal speed |
| **Evidence-backed bare-edge resolution** | Prevents accidental cross-repo name collisions by requiring import/file evidence |
| **Two-phase search (FTS5 → LIKE)** | Fast tokenized search on modern schemas, graceful fallback on older databases |
| **Language-family grouping** | Reduces false positives when calls cross JS/TS/TSX boundaries |
| **Daemon incremental updates** | Maintains fresh graphs without full re-parsing, critical for CI efficiency |

## Summary

- **Code Review Graph architecture** centers on a four-layer design: **persistence**, **parsing**, **language resolution**, and **service APIs**
- The `GraphStore` in [`graph.py`](https://github.com/tirth8205/code-review-graph/blob/main/graph.py) provides transactional SQLite storage with an in-memory NetworkX cache for performance
- Language-specific resolvers in [`python_resolver.py`](https://github.com/tirth8205/code-review-graph/blob/main/python_resolver.py), [`spring_resolver.py`](https://github.com/tirth8205/code-review-graph/blob/main/spring_resolver.py), and related files extract semantically meaningful relationships
- The service layer exposes functionality through CLI, daemon, search, impact analysis, and visualization modules
- Post-import resolution and evidence-backed edge linking ensure graph accuracy despite ambiguous source references

## Frequently Asked Questions

### What database does Code Review Graph use?

Code Review Graph uses **SQLite** as its primary persistence layer. This choice guarantees durability and fast random access without requiring external database infrastructure. The `GraphStore` class in [`graph.py`](https://github.com/tirth8205/code-review-graph/blob/main/graph.py) manages all database operations, while a cached **NetworkX** graph provides O(1) edge lookups for in-memory traversals.

### How does Code Review Graph handle multiple programming languages?

Each supported language has a dedicated **resolver module** that understands its grammar and module system. The [`parser.py`](https://github.com/tirth8205/code-review-graph/blob/main/parser.py) orchestrator selects the appropriate resolver based on file extension—routing `.py` files to [`python_resolver.py`](https://github.com/tirth8205/code-review-graph/blob/main/python_resolver.py), Spring Java files to [`spring_resolver.py`](https://github.com/tirth8205/code-review-graph/blob/main/spring_resolver.py), and so on. Resolvers return standardized `NodeInfo` and `EdgeInfo` objects that the graph stores uniformly.

### What is "impact radius" in Code Review Graph?

**Impact radius** is a graph traversal feature that identifies all code potentially affected by a change. The `get_impact_radius` function (available through CLI as `crg impact`) walks outgoing or incoming edges from a starting symbol to a configurable depth, using edge-type and direction constants from [`constants.py`](https://github.com/tirth8205/code-review-graph/blob/main/constants.py). This enables automated risk assessment during code review.

### How does Code Review Graph stay updated with code changes?

The **daemon mode** ([`daemon.py`](https://github.com/tirth8205/code-review-graph/blob/main/daemon.py)) runs as a background process that watches the repository filesystem, detects modified files, and re-indexes only those changes. This incremental approach keeps the graph current without the latency of full re-parsing, making it suitable for continuous integration pipelines and active development workflows.