How the Incremental Update Engine in code-review-graph Detects and Processes File Changes

The incremental update engine in code-review-graph detects file changes by comparing the current Git HEAD against a stored SHA, expands the change set through graph-based impact analysis, and selectively re-parses only affected files while skipping unchanged content via SHA-256 hashing.

The incremental update engine lives in code_review_graph/incremental.py and is designed to keep the code-review graph synchronized with repository changes with minimal overhead. Rather than rebuilding the entire graph on every update, it employs a three-stage pipeline: change detection, impact analysis, and selective parsing with graph mutation. This article examines each stage with direct references to the source implementation in the tirth8205/code-review-graph repository.


Change Detection: Finding What Changed

The engine starts by identifying exactly which files the version control system reports as modified since the last successful build.

Resolving the Diff Base

The function resolve_incremental_base() (lines 644-667) determines the starting point for comparison:


# From code_review_graph/incremental.py

def resolve_incremental_base(repo_root: Path, store: GraphStore) -> Optional[str]:
    """
    Reads the last stored Git SHA (git_head_sha) from the GraphStore.
    Returns None if the commit no longer exists or if no base is stored,
    triggering a full rebuild instead.
    """
    stored_sha = store.get_metadata("git_head_sha")
    if stored_sha and _commit_object_exists(repo_root, stored_sha):
        return stored_sha
    return None  # Forces full rebuild

The stored SHA is validated against _SAFE_GIT_REF (lines 775-784), a regex that prevents injection attacks and malformed references before any Git command execution.

Running the VCS Diff

The get_changed_files() function (lines 670-708) delegates to Git or SVN. For Git repositories, it executes:

git diff --name-status -z <base> --

The -z flag produces NUL-separated output, which _decode_name_status_paths() (lines 780-802) parses to handle:

  • Added files (status A)
  • Modified files (status M)
  • Deleted files (status D)
  • Renamed/copied files (status R or C) — which emit two paths (old and new)

# From _decode_name_status_paths (lines 780-802)

def _decode_name_status_paths(raw: bytes) -> List[str]:
    """
    Decodes NUL-separated git diff --name-status -z output.
    Handles rename pairs by extracting both source and destination.
    """
    parts = raw.split(b'\x00')
    paths = []
    i = 0
    while i < len(parts) - 1:
        status = parts[i].decode()
        if status.startswith('R') or status.startswith('C'):
            # Rename/copy: two paths follow

            paths.extend([parts[i+1].decode(), parts[i+2].decode()])
            i += 3
        else:
            paths.append(parts[i+1].decode())
            i += 2
    return list(dict.fromkeys(paths))  # Deduplicate while preserving order

The result is a deduplicated list changed_files containing every path that requires attention, including both sides of renames so stale entries can be purged later.


Impact Analysis: Expanding to Dependents

Detecting direct changes is insufficient — the engine must also identify files that depend on changed files through imports, calls, or inheritance.

Reconciling Stale Files

The _reconcile_stale_files() function (lines 909-941) compares the repository inventory against the graph's stored file list. Files present in the graph but missing from disk are flagged for removal.

Finding Dependent Files

The find_dependents() function (lines 986-1024) performs a bounded graph walk:


# From code_review_graph/incremental.py

def find_dependents(store: GraphStore, changed_file: str, max_hops: int = _MAX_DEPENDENT_HOPS) -> Set[str]:
    """
    Walks the graph upward from changed_file to find files that reference it.
    Default max_hops is 2, preventing explosion on highly connected graphs.
    """
    dependents = set()
    current_layer = {changed_file}
    for hop in range(max_hops):
        next_layer = set()
        for f in current_layer:
            next_layer.update(_single_hop_dependents(store, f))
        dependents.update(next_layer)
        current_layer = next_layer - dependents
    return dependents

The combined set all_files = changed_files ∪ dependent_files is then filtered through _should_ignore() to exclude binary files and paths matching project ignore patterns (lines 1232-1246).


Selective Parsing and Graph Mutation

With the minimal file set identified, the engine applies multiple optimizations to avoid unnecessary work.

Content Hash Deduplication

Before parsing any file, the engine computes its SHA-256 hash and compares against stored hashes (lines 1245-1251):


# Inside incremental_update (lines 1245-1251)

file_hash = hashlib.sha256(file_path.read_bytes()).hexdigest()
stored_hash = store.get_file_hash(rel_path)
if file_hash == stored_hash:
    continue  # Skip: file unchanged despite VCS report

This catches cases where timestamps differ but content is identical — common after Git operations like rebases or branch switches.

Parallel vs. Serial Parsing

The engine automatically selects execution strategy (lines 1260-1298):

  • Serial parsing — when CRG_SERIAL_PARSE=1 or file count is small
  • Parallel parsing — using _make_executor() with kind selected by _select_executor_kind() (processes for CPU-bound work, threads for I/O-bound)

Worker Function

Each file is processed by _parse_single_file() (lines 1028-1050):

def _parse_single_file(repo_root: Path, rel_path: str, file_hash: str) -> ParsedResult:
    """
    Thread-local parser creation, file reading, and AST analysis.
    Returns nodes, edges, and computed hash for persistence.
    """
    full_path = repo_root / rel_path
    source_bytes = full_path.read_bytes()
    
    # Lazy thread-local parser initialization

    if not hasattr(_local, 'parser'):
        _local.parser = CodeParser()
    
    nodes, edges = _local.parser.parse(rel_path, source_bytes)
    return ParsedResult(nodes=nodes, edges=edges, file_hash=file_hash)

Database Operations

Results are persisted through store.store_file_nodes_edges() (lines 1268-1285), with:

  • Single transaction commit after all files complete
  • Permanent removal of deleted files via store.remove_files_permanently() (lines 1299-1302)
  • Metadata refresh for branch, head_sha, and last_updated timestamp (lines 1304-1307)

Conditional Resolver Execution

Language-specific post-processing runs only when needed (lines 1310-1332):


# Resolver gating example

if any(f.endswith('.py') for f in all_files):
    run_python_import_resolver(store)
if any('spring' in f.lower() for f in all_files):
    run_spring_di_resolver(store)

Running Incremental Updates

CLI Usage


# From repository root

crg update          # Uses stored HEAD SHA as diff base

# Force full rebuild

crg update --full

Programmatic API

from code_review_graph.incremental import incremental_update, get_changed_files
from code_review_graph.graph import GraphStore
from pathlib import Path

repo_root = Path(".")
store = GraphStore(repo_root / ".code-review-graph/graph.db")

# Basic incremental update

result = incremental_update(repo_root, store)
print(f"Updated {result['files_updated']} files")

# Manual change detection for custom logic

base = resolve_incremental_base(repo_root, store) or "HEAD~1"
changed = get_changed_files(repo_root, base)
print(f"VCS detected: {len(changed)} changed files")

Expanding Impact Sets Manually

from code_review_graph.incremental import find_dependents

changed = ["src/services/user.ts", "src/models/order.py"]

for rel_path in changed:
    dependents = find_dependents(store, rel_path)
    impact_set = {rel_path} | dependents
    print(f"{rel_path}: {len(dependents)} dependents, {len(impact_set)} total to process")

Summary

  • Diff base resolution (resolve_incremental_base, lines 644-667) validates stored SHAs and falls back to full rebuilds when history is unreachable
  • VCS integration (get_changed_files, lines 670-708) handles Git and SVN with NUL-separated output parsing for rename detection
  • Impact expansion (find_dependents, lines 986-1024) walks up to _MAX_DEPENDENT_HOPS (default 2) to find referencing files
  • Content hashing prevents re-parsing of identical files even when VCS timestamps differ
  • Adaptive execution chooses serial or parallel parsing based on workload size and environment variables
  • Conditional resolvers avoid running language-specific analysis when no relevant files changed

The incremental update engine in code-review-graph achieves near-instant graph synchronization on large codebases by treating the graph itself as a cache and minimizing all operations — I/O, parsing, and database writes — to the theoretical minimum required for correctness.


Frequently Asked Questions

How does the incremental update engine handle Git renames?

The engine detects renames through git diff --name-status -z, which emits two paths for R (rename) and C (copy) statuses. The _decode_name_status_paths() function (lines 780-802) extracts both the old and new paths, ensuring the old file's nodes are removed from the graph while the new file is parsed fresh. This prevents orphaned nodes and broken references according to the tirth8205/code-review-graph source code.

What triggers a full rebuild instead of an incremental update?

A full rebuild occurs when resolve_incremental_base() returns None, which happens if: (1) no git_head_sha is stored in the database, (2) the stored SHA no longer exists in the repository (e.g., after history rewriting), or (3) _commit_object_exists() fails to locate the commit object. The safety check against _SAFE_GIT_REF (lines 775-784) also rejects malformed references that could indicate corruption or tampering.

Why does the engine hash file contents after already running Git diff?

Git's diff relies on metadata like timestamps and mode bits, which can change without content changes — common after git rebase, git cherry-pick, or branch switches. The SHA-256 content hash check (lines 1245-1251) provides a cryptographic guarantee of content equality, skipping parsing for files that are bit-for-bit identical to their stored versions. This second-layer deduplication significantly reduces parsing overhead in active development workflows.

How does parallel parsing avoid race conditions?

The _parse_single_file() function (lines 1028-1050) uses thread-local storage via _local to create isolated CodeParser instances per thread. Database writes are batched and committed in a single main-thread transaction after all workers complete, preventing concurrent access to the SQLite store. The executor kind (process vs. thread) is selected by _select_executor_kind() based on whether the workload is CPU-bound or I/O-bound.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →