# Performance Considerations for code‑review‑graph: 8 Optimizations That Keep Large‑Scale Queries Fast

> Discover 8 performance optimizations for code-review-graph handling millions of nodes. Learn about WAL mode SQLite, indexing, batch transactions, and more to keep large scale queries fast.

- Repository: [Tirth Kanani/code-review-graph](https://github.com/tirth8205/code-review-graph)
- Tags: performance
- Published: 2026-08-17

---

**The code‑review‑graph repository optimizes performance through WAL‑mode SQLite, strategic indexing, batch transactions, lazy NetworkX caching, and hard‑capped graph traversals to handle millions of nodes without blocking or exponential slowdowns.**

The **code‑review‑graph** project stores a full‑text, cross‑language code‑structure graph in an embedded SQLite database and exposes powerful query tools for impact analysis, flow tracing, and search. Because these operations can touch millions of nodes and edges, the implementation makes deliberate architectural trade‑offs between memory usage and I/O latency. This article examines the eight key performance optimizations baked into the codebase, with references to specific source files and runnable patterns.

## Fast Concurrent Reads with WAL Mode

SQLite's default rollback journal blocks readers during writes. The codebase avoids this entirely by opening every database connection in **Write‑Ahead Logging (WAL) mode**.

In `code_review_graph/graph.py:191`, the connection initialization executes:

```python
PRAGMA journal_mode=WAL

```

This allows readers to proceed without blocking on writers, and writers to commit without waiting for readers. For a tool that ingests files continuously while serving interactive queries, this eliminates lock contention as a bottleneck.

## Efficient Indexing on Hot Columns

The `nodes` and `edges` tables carry indexes on exactly the columns used in production query paths. As defined in `graph.py:112‑121`, these include:

- `qualified_name` — for precise symbol lookup
- `file_path` — for file‑scoped queries and incremental updates
- `kind` — for node type filtering
- `source_qualified` and `target_qualified` — for edge traversal

These **covering indexes** reduce lookups to index‑only scans in the common case, minimizing page reads from disk.

## Batch Writes to Eliminate Transaction Churn

Per‑statement commits would cripple ingestion throughput. The `store_file_nodes_edges` and `store_file_batch` methods in `graph.py:159‑170` wrap all mutations for a single file in one transaction:

```python
with self._transaction() as cur:
    # bulk insert nodes

    # bulk insert edges

    # commit once

```

By using `BEGIN IMMEDIATE`, the code acquires the write lock upfront and holds it for the entire batch, reducing lock acquisition overhead and fsync frequency.

## Lazy In‑Memory NetworkX Caching

Graph algorithms run faster in memory, but rebuilding a `networkx.DiGraph` on every query is wasteful. The implementation in `graph.py:208‑224` maintains a **`self._nxg_cache`** that is:

- Built lazily on first read‑only operation
- Returned directly for subsequent traversals
- Invalidated automatically after any write via `_invalidate_cache`

This trades a modest memory footprint (one Python object per node and edge) for sub‑millisecond response times on complex traversals.

## Controlled Blast‑Radius Traversal

Unbounded BFS on code graphs hits hub functions (like logging utilities) and explodes combinatorially. The `get_impact_radius` method enforces three hard limits to guarantee **deterministic response times**:

| Limit | Purpose | Default |
|-------|---------|---------|
| `MAX_IMPACT_DEPTH` | Cap BFS depth | Configurable at call site |
| `MAX_IMPACT_NODES` | Cap total nodes visited | Via `max_results` parameter |
| `CRG_MAX_TRANSITIVE_FRONTIER` | Cap per‑hop frontier expansion | Prevents hub‑induced explosion |

These limits are implemented in `graph.py:150‑166` and exposed through `tools/query.py#get_impact_radius`. Callers can tighten or relax them based on their latency requirements.

## Bare‑Name Resolution with Evidence‑Backed Caches

Resolving ambiguous `CALLS` or `TESTED_BY` edges requires matching bare identifiers against candidate definitions. Rather than issuing repeated SQL queries, the resolver in `graph.py:941‑979` builds **in‑memory maps**:

- Candidate nodes by simple name
- Imported file scopes for each context
- Namespace evidence for disambiguation

This single‑pass pre‑loading replaces N round‑trips with one bulk fetch and Python‑side matching, cutting resolution time by orders of magnitude on large files.

## Full‑Text Search with FTS5 Fallback

The `search_nodes` method in `graph.py:1034‑1055` implements a **fast‑path/slow‑path** pattern:

```python

# Attempt FTS5 virtual table first

matches = self._fts_search(query, limit)
if not matches:
    # Graceful fallback to LIKE scan

    matches = self._like_search(query, limit)

```

The `nodes_fts` table uses SQLite's **FTS5** extension for tokenized, ranked matching. If FTS5 is unavailable or returns no hits, the code falls back to a cheap substring scan rather than failing or performing a full table scan.

## Incremental Processing for Large Repositories

Full re‑ingestion of a monorepo is prohibitive. The [`incremental.py`](https://github.com/tirth8205/code-review-graph/blob/main/incremental.py) module (referenced but not detailed in source) identifies changed files through content hashing, then invokes `GraphStore.store_file_batch` to rewrite only affected graph regions.

This **delta‑oriented design** keeps daily synchronization operations bounded by edit size rather than repository size.

## Thread‑Safe Connection Handling

SQLite connections are not thread‑safe by default. The codebase in `graph.py:190‑197` creates connections with:

```python
sqlite3.connect(db_path, check_same_thread=False)

```

A per‑instance lock protects the NetworkX cache during concurrent reads, allowing safe multi‑threaded access to the `GraphStore` from web servers or parallel pipelines.

## Practical Usage Pattern

The following demonstrates the performance‑aware API in action:

```python
from code_review_graph.graph import GraphStore
from pathlib import Path

# Open with all optimizations active

store = GraphStore(Path("~/.crg/graph.db").expanduser())

# Ingest with single transaction

store.store_file_nodes_edges(
    file_path="src/example.py",
    nodes=parsed_nodes,
    edges=parsed_edges,
    fhash="d41d8cd98f00b204e9800998ecf8427e",
)

# Capped impact analysis

impact = store.get_impact_radius(
    changed_files=["src/example.py"],
    max_depth=3,
    max_results=200,
)

# Fast full‑text lookup

matches = store.search_nodes("authentication token", limit=10)

# Bulk resolution without repeated queries

resolved = store.resolve_bare_call_targets()

```

All heavy lifting—transactions, caching, BFS limits—remains internal to `GraphStore`, presenting a simple responsive surface.

## Summary

- **WAL mode** enables concurrent reads without writer blocking
- **Strategic indexes** on `qualified_name`, `file_path`, and edge endpoints optimize lookup paths
- **Batch transactions** in `store_file_nodes_edges` minimize commit overhead
- **Lazy NetworkX cache** amortizes graph construction cost across queries
- **Hard caps** on depth, node count, and frontier size prevent traversal explosions
- **In‑memory resolution maps** eliminate SQL round‑trips during bare‑name matching
- **FTS5‑first search** with `LIKE` fallback balances speed and compatibility
- **Incremental updates** bound processing cost to changes rather than total repository size

## Frequently Asked Questions

### Does code‑review‑graph support concurrent queries from multiple threads?

Yes. The `GraphStore` opens SQLite with `check_same_thread=False` and protects the NetworkX cache with an instance‑level lock. Multiple threads can query simultaneously; writers serialize through SQLite's WAL mode without blocking readers.

### What happens if a blast‑radius query hits a widely used utility function?

The `get_impact_radius` implementation enforces three independent limits: maximum BFS depth, total nodes visited, and per‑hop frontier size. These caps terminate traversal early on hub functions, returning a bounded partial result rather than hanging or exhausting memory.

### How does the repository handle full‑text search without external dependencies?

The search layer uses SQLite's built‑in **FTS5** extension when available, falling back to a `LIKE`-based scan if FTS5 is unavailable or yields no matches. No external search engine is required, keeping deployment self‑contained.

### Can I use code‑review‑graph on a repository too large to fit in memory?

Yes. The core storage is SQLite on disk; only the working set needed for a specific query loads into memory. The NetworkX cache is optional and can be disabled or size‑limited. Incremental ingestion ensures daily updates remain proportional to edits, not total code volume.