Incremental Update Performance for Large Projects Like Django: A Deep Dive into code-review-graph

code-review-graph completes incremental updates on large codebases like Django in approximately 2.5 seconds for a typical two-file edit, with most of that time spent on Python process startup rather than actual re-indexing.

The code-review-graph tool from tirth8205/code-review-graph solves a critical problem for developers working with massive Python repositories: how to keep an up-to-date knowledge graph without re-parsing thousands of files on every change. This article examines how its incremental update system performs on enterprise-scale projects, using Django's ~3,000-file codebase as the benchmark.

How Incremental Updates Work in code-review-graph

The incremental pipeline is orchestrated in [code_review_graph/incremental.py](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/incremental.py#L61-L70) through the incremental_update function. The process follows four distinct phases:

  1. Detect changed files — Uses git diff or svn status to identify modifications since the last update
  2. Remove stale entries — Calls _reconcile_stale_files to purge outdated graph nodes
  3. Find dependents — Traces import and call edges via find_dependents to locate downstream files affected by the changes
  4. Re-parse selectively — Only processes files whose SHA-256 hash actually differs from the cached version

This minimal-reprocessing strategy ensures that unchanged files are never touched, regardless of project size.

Measured Performance on Django

The repository's README documents concrete benchmarks for a shallow clone of Django containing roughly 3,000 Python files:

Scenario Latency Breakdown
Two-file edit ~2.5 seconds ~1.4s process startup + ~1.1s actual work
No-op update ~1.4 seconds Startup cost only

These measurements were obtained using the tool's built-in timing capabilities and external time command. The dominant cost is Python process startup, not graph computation — making subsequent no-op updates essentially free once the interpreter is warm.

Triggering Incremental Updates

CLI: Watch Mode and Manual Updates

The most common workflow uses automatic watch mode:


# Automatically update on every file save

code-review-graph watch

For manual invocation with timing visibility:


# Trigger update and measure wall-clock time

time code-review-graph update --brief

The --brief flag displays the token-savings panel while time reveals the actual latency.

Programmatic Python API

For custom tooling or CI integration:

from pathlib import Path
from code_review_graph.graph import GraphStore
from code_review_graph.incremental import incremental_update, full_build

repo_root = Path("/path/to/django")
store = GraphStore(repo_root / ".code-review-graph" / "graph.db")

# One-time full build

full_build(repo_root, store)

# Subsequent incremental updates after file changes

result = incremental_update(repo_root, store)
print(f"Files reparsed: {result['files_updated']}")
print(f"Total nodes now: {store.get_stats().total_nodes}")

Measuring Latency Programmatically

To capture precise timing within Python:

import time

start = time.time()
incremental_update(repo_root, store)
elapsed = time.time() - start
print(f"Incremental update took {elapsed:.2f}s")

Key Implementation Details

Diff Base Resolution

Before detecting changes, incremental_update calls resolve_incremental_base to determine the appropriate comparison point. This handles various Git states cleanly, including detached HEAD and shallow clones.

Parallel Execution

The parsing phase uses a parallel executor to process multiple changed files concurrently. This is particularly effective when edits cascade through many dependents.

Language-Specific Resolvers

After the core graph update completes, language-specific resolvers run only when their relevant files have changed. This avoids redundant overhead for unmodified language domains.

Core Source Files

Understanding these files provides deeper insight into the performance characteristics:

File Purpose
code_review_graph/incremental.py Main implementation: incremental_update, change detection, dependent discovery
code_review_graph/graph.py GraphStore class for persistence and statistics retrieval
README.md (Incremental update latency section) Official benchmarks for Django-scale repositories
tests/test_incremental.py Verification suite for incremental correctness
diagrams/diagram4_incremental_update.png Visual flow diagram of the update pipeline

Summary

  • incremental_update in code_review_graph/incremental.py drives all incremental processing through a four-phase pipeline
  • SHA-256 hash comparison eliminates redundant parsing of unchanged files
  • Django benchmark: ~2.5 seconds for typical edits, with ~1.4 seconds attributable to Python startup overhead
  • No-op updates complete in ~1.4 seconds (startup cost only)
  • Watch mode and manual CLI both leverage the same optimized code path
  • Parallel parsing and conditional resolver execution minimize active computation time

Frequently Asked Questions

How does code-review-graph determine which files need re-parsing?

The tool computes SHA-256 hashes for all source files and compares them against cached values from the previous run. Only files with hash mismatches are re-parsed, as implemented in the incremental pipeline of code_review_graph/incremental.py. This hash-based approach is deterministic and ignores file timestamps, avoiding false positives from Git operations or filesystem quirks.

What happens to dependent files when I edit a single module?

The find_dependents function traces import and call edges through the existing graph to identify downstream files affected by your changes. These dependents are added to the re-parse queue even if their own source code hasn't changed, ensuring the knowledge graph remains consistent. The reconciliation phase _reconcile_stale_files removes outdated entries before the new parsing begins.

Can I use incremental updates in CI/CD pipelines?

Yes, the programmatic API supports this workflow. Initialize a GraphStore pointing to a persistent location, run full_build once to bootstrap, then call incremental_update on subsequent pipeline runs. Store the graph database as a build artifact or cache it between runs to benefit from incremental speedups. The result['files_updated'] return value helps audit what changed between builds.

Why is process startup the dominant cost rather than parsing?

Python interpreter initialization dominates because the actual graph operations are highly optimized: minimal file I/O, selective parsing, and in-memory graph mutations. For the Django benchmark, only two files required re-parsing — a trivial amount of work compared to launching the Python runtime. Warm-start scenarios (long-running watch mode) eliminate this cost entirely, reducing effective latency to approximately one second.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →