How the Incremental Update Engine in code-review-graph Detects and Processes File Changes
The incremental update engine in code-review-graph detects file changes by comparing the current Git HEAD against a stored SHA, expands the change set through graph-based impact analysis, and selectively re-parses only affected files while skipping unchanged content via SHA-256 hashing.
The incremental update engine lives in code_review_graph/incremental.py and is designed to keep the code-review graph synchronized with repository changes with minimal overhead. Rather than rebuilding the entire graph on every update, it employs a three-stage pipeline: change detection, impact analysis, and selective parsing with graph mutation. This article examines each stage with direct references to the source implementation in the tirth8205/code-review-graph repository.
Change Detection: Finding What Changed
The engine starts by identifying exactly which files the version control system reports as modified since the last successful build.
Resolving the Diff Base
The function resolve_incremental_base() (lines 644-667) determines the starting point for comparison:
# From code_review_graph/incremental.py
def resolve_incremental_base(repo_root: Path, store: GraphStore) -> Optional[str]:
"""
Reads the last stored Git SHA (git_head_sha) from the GraphStore.
Returns None if the commit no longer exists or if no base is stored,
triggering a full rebuild instead.
"""
stored_sha = store.get_metadata("git_head_sha")
if stored_sha and _commit_object_exists(repo_root, stored_sha):
return stored_sha
return None # Forces full rebuild
The stored SHA is validated against _SAFE_GIT_REF (lines 775-784), a regex that prevents injection attacks and malformed references before any Git command execution.
Running the VCS Diff
The get_changed_files() function (lines 670-708) delegates to Git or SVN. For Git repositories, it executes:
git diff --name-status -z <base> --
The -z flag produces NUL-separated output, which _decode_name_status_paths() (lines 780-802) parses to handle:
- Added files (status
A) - Modified files (status
M) - Deleted files (status
D) - Renamed/copied files (status
RorC) — which emit two paths (old and new)
# From _decode_name_status_paths (lines 780-802)
def _decode_name_status_paths(raw: bytes) -> List[str]:
"""
Decodes NUL-separated git diff --name-status -z output.
Handles rename pairs by extracting both source and destination.
"""
parts = raw.split(b'\x00')
paths = []
i = 0
while i < len(parts) - 1:
status = parts[i].decode()
if status.startswith('R') or status.startswith('C'):
# Rename/copy: two paths follow
paths.extend([parts[i+1].decode(), parts[i+2].decode()])
i += 3
else:
paths.append(parts[i+1].decode())
i += 2
return list(dict.fromkeys(paths)) # Deduplicate while preserving order
The result is a deduplicated list changed_files containing every path that requires attention, including both sides of renames so stale entries can be purged later.
Impact Analysis: Expanding to Dependents
Detecting direct changes is insufficient — the engine must also identify files that depend on changed files through imports, calls, or inheritance.
Reconciling Stale Files
The _reconcile_stale_files() function (lines 909-941) compares the repository inventory against the graph's stored file list. Files present in the graph but missing from disk are flagged for removal.
Finding Dependent Files
The find_dependents() function (lines 986-1024) performs a bounded graph walk:
# From code_review_graph/incremental.py
def find_dependents(store: GraphStore, changed_file: str, max_hops: int = _MAX_DEPENDENT_HOPS) -> Set[str]:
"""
Walks the graph upward from changed_file to find files that reference it.
Default max_hops is 2, preventing explosion on highly connected graphs.
"""
dependents = set()
current_layer = {changed_file}
for hop in range(max_hops):
next_layer = set()
for f in current_layer:
next_layer.update(_single_hop_dependents(store, f))
dependents.update(next_layer)
current_layer = next_layer - dependents
return dependents
The combined set all_files = changed_files ∪ dependent_files is then filtered through _should_ignore() to exclude binary files and paths matching project ignore patterns (lines 1232-1246).
Selective Parsing and Graph Mutation
With the minimal file set identified, the engine applies multiple optimizations to avoid unnecessary work.
Content Hash Deduplication
Before parsing any file, the engine computes its SHA-256 hash and compares against stored hashes (lines 1245-1251):
# Inside incremental_update (lines 1245-1251)
file_hash = hashlib.sha256(file_path.read_bytes()).hexdigest()
stored_hash = store.get_file_hash(rel_path)
if file_hash == stored_hash:
continue # Skip: file unchanged despite VCS report
This catches cases where timestamps differ but content is identical — common after Git operations like rebases or branch switches.
Parallel vs. Serial Parsing
The engine automatically selects execution strategy (lines 1260-1298):
- Serial parsing — when
CRG_SERIAL_PARSE=1or file count is small - Parallel parsing — using
_make_executor()with kind selected by_select_executor_kind()(processes for CPU-bound work, threads for I/O-bound)
Worker Function
Each file is processed by _parse_single_file() (lines 1028-1050):
def _parse_single_file(repo_root: Path, rel_path: str, file_hash: str) -> ParsedResult:
"""
Thread-local parser creation, file reading, and AST analysis.
Returns nodes, edges, and computed hash for persistence.
"""
full_path = repo_root / rel_path
source_bytes = full_path.read_bytes()
# Lazy thread-local parser initialization
if not hasattr(_local, 'parser'):
_local.parser = CodeParser()
nodes, edges = _local.parser.parse(rel_path, source_bytes)
return ParsedResult(nodes=nodes, edges=edges, file_hash=file_hash)
Database Operations
Results are persisted through store.store_file_nodes_edges() (lines 1268-1285), with:
- Single transaction commit after all files complete
- Permanent removal of deleted files via
store.remove_files_permanently()(lines 1299-1302) - Metadata refresh for
branch,head_sha, andlast_updatedtimestamp (lines 1304-1307)
Conditional Resolver Execution
Language-specific post-processing runs only when needed (lines 1310-1332):
# Resolver gating example
if any(f.endswith('.py') for f in all_files):
run_python_import_resolver(store)
if any('spring' in f.lower() for f in all_files):
run_spring_di_resolver(store)
Running Incremental Updates
CLI Usage
# From repository root
crg update # Uses stored HEAD SHA as diff base
# Force full rebuild
crg update --full
Programmatic API
from code_review_graph.incremental import incremental_update, get_changed_files
from code_review_graph.graph import GraphStore
from pathlib import Path
repo_root = Path(".")
store = GraphStore(repo_root / ".code-review-graph/graph.db")
# Basic incremental update
result = incremental_update(repo_root, store)
print(f"Updated {result['files_updated']} files")
# Manual change detection for custom logic
base = resolve_incremental_base(repo_root, store) or "HEAD~1"
changed = get_changed_files(repo_root, base)
print(f"VCS detected: {len(changed)} changed files")
Expanding Impact Sets Manually
from code_review_graph.incremental import find_dependents
changed = ["src/services/user.ts", "src/models/order.py"]
for rel_path in changed:
dependents = find_dependents(store, rel_path)
impact_set = {rel_path} | dependents
print(f"{rel_path}: {len(dependents)} dependents, {len(impact_set)} total to process")
Summary
- Diff base resolution (
resolve_incremental_base, lines 644-667) validates stored SHAs and falls back to full rebuilds when history is unreachable - VCS integration (
get_changed_files, lines 670-708) handles Git and SVN with NUL-separated output parsing for rename detection - Impact expansion (
find_dependents, lines 986-1024) walks up to_MAX_DEPENDENT_HOPS(default 2) to find referencing files - Content hashing prevents re-parsing of identical files even when VCS timestamps differ
- Adaptive execution chooses serial or parallel parsing based on workload size and environment variables
- Conditional resolvers avoid running language-specific analysis when no relevant files changed
The incremental update engine in code-review-graph achieves near-instant graph synchronization on large codebases by treating the graph itself as a cache and minimizing all operations — I/O, parsing, and database writes — to the theoretical minimum required for correctness.
Frequently Asked Questions
How does the incremental update engine handle Git renames?
The engine detects renames through git diff --name-status -z, which emits two paths for R (rename) and C (copy) statuses. The _decode_name_status_paths() function (lines 780-802) extracts both the old and new paths, ensuring the old file's nodes are removed from the graph while the new file is parsed fresh. This prevents orphaned nodes and broken references according to the tirth8205/code-review-graph source code.
What triggers a full rebuild instead of an incremental update?
A full rebuild occurs when resolve_incremental_base() returns None, which happens if: (1) no git_head_sha is stored in the database, (2) the stored SHA no longer exists in the repository (e.g., after history rewriting), or (3) _commit_object_exists() fails to locate the commit object. The safety check against _SAFE_GIT_REF (lines 775-784) also rejects malformed references that could indicate corruption or tampering.
Why does the engine hash file contents after already running Git diff?
Git's diff relies on metadata like timestamps and mode bits, which can change without content changes — common after git rebase, git cherry-pick, or branch switches. The SHA-256 content hash check (lines 1245-1251) provides a cryptographic guarantee of content equality, skipping parsing for files that are bit-for-bit identical to their stored versions. This second-layer deduplication significantly reduces parsing overhead in active development workflows.
How does parallel parsing avoid race conditions?
The _parse_single_file() function (lines 1028-1050) uses thread-local storage via _local to create isolated CodeParser instances per thread. Database writes are batched and committed in a single main-thread transaction after all workers complete, preventing concurrent access to the SQLite store. The executor kind (process vs. thread) is selected by _select_executor_kind() based on whether the workload is CPU-bound or I/O-bound.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →