# How the Incremental Update Engine in code-review-graph Detects and Processes File Changes

> Learn how the incremental update engine in code-review-graph efficiently detects and processes file changes by comparing SHAs, performing graph-based analysis, and re-parsing only affected files.

- Repository: [Tirth Kanani/code-review-graph](https://github.com/tirth8205/code-review-graph)
- Tags: internals
- Published: 2026-08-11

---

**The incremental update engine in code-review-graph detects file changes by comparing the current Git HEAD against a stored SHA, expands the change set through graph-based impact analysis, and selectively re-parses only affected files while skipping unchanged content via SHA-256 hashing.**

The **incremental update engine** lives in [`code_review_graph/incremental.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/incremental.py) and is designed to keep the code-review graph synchronized with repository changes with minimal overhead. Rather than rebuilding the entire graph on every update, it employs a three-stage pipeline: **change detection**, **impact analysis**, and **selective parsing with graph mutation**. This article examines each stage with direct references to the source implementation in the tirth8205/code-review-graph repository.

---

## Change Detection: Finding What Changed

The engine starts by identifying exactly which files the version control system reports as modified since the last successful build.

### Resolving the Diff Base

The function `resolve_incremental_base()` (lines 644-667) determines the starting point for comparison:

```python

# From code_review_graph/incremental.py

def resolve_incremental_base(repo_root: Path, store: GraphStore) -> Optional[str]:
    """
    Reads the last stored Git SHA (git_head_sha) from the GraphStore.
    Returns None if the commit no longer exists or if no base is stored,
    triggering a full rebuild instead.
    """
    stored_sha = store.get_metadata("git_head_sha")
    if stored_sha and _commit_object_exists(repo_root, stored_sha):
        return stored_sha
    return None  # Forces full rebuild

```

The stored SHA is validated against `_SAFE_GIT_REF` (lines 775-784), a regex that prevents injection attacks and malformed references before any Git command execution.

### Running the VCS Diff

The `get_changed_files()` function (lines 670-708) delegates to Git or SVN. For Git repositories, it executes:

```bash
git diff --name-status -z <base> --

```

The `-z` flag produces NUL-separated output, which `_decode_name_status_paths()` (lines 780-802) parses to handle:

- **Added files** (status `A`)
- **Modified files** (status `M`)
- **Deleted files** (status `D`)
- **Renamed/copied files** (status `R` or `C`) — which emit **two paths** (old and new)

```python

# From _decode_name_status_paths (lines 780-802)

def _decode_name_status_paths(raw: bytes) -> List[str]:
    """
    Decodes NUL-separated git diff --name-status -z output.
    Handles rename pairs by extracting both source and destination.
    """
    parts = raw.split(b'\x00')
    paths = []
    i = 0
    while i < len(parts) - 1:
        status = parts[i].decode()
        if status.startswith('R') or status.startswith('C'):
            # Rename/copy: two paths follow

            paths.extend([parts[i+1].decode(), parts[i+2].decode()])
            i += 3
        else:
            paths.append(parts[i+1].decode())
            i += 2
    return list(dict.fromkeys(paths))  # Deduplicate while preserving order

```

The result is a deduplicated list `changed_files` containing every path that requires attention, including both sides of renames so stale entries can be purged later.

---

## Impact Analysis: Expanding to Dependents

Detecting direct changes is insufficient — the engine must also identify files that **depend** on changed files through imports, calls, or inheritance.

### Reconciling Stale Files

The `_reconcile_stale_files()` function (lines 909-941) compares the repository inventory against the graph's stored file list. Files present in the graph but missing from disk are flagged for removal.

### Finding Dependent Files

The `find_dependents()` function (lines 986-1024) performs a bounded graph walk:

```python

# From code_review_graph/incremental.py

def find_dependents(store: GraphStore, changed_file: str, max_hops: int = _MAX_DEPENDENT_HOPS) -> Set[str]:
    """
    Walks the graph upward from changed_file to find files that reference it.
    Default max_hops is 2, preventing explosion on highly connected graphs.
    """
    dependents = set()
    current_layer = {changed_file}
    for hop in range(max_hops):
        next_layer = set()
        for f in current_layer:
            next_layer.update(_single_hop_dependents(store, f))
        dependents.update(next_layer)
        current_layer = next_layer - dependents
    return dependents

```

The combined set `all_files = changed_files ∪ dependent_files` is then filtered through `_should_ignore()` to exclude binary files and paths matching project ignore patterns (lines 1232-1246).

---

## Selective Parsing and Graph Mutation

With the minimal file set identified, the engine applies multiple optimizations to avoid unnecessary work.

### Content Hash Deduplication

Before parsing any file, the engine computes its **SHA-256 hash** and compares against stored hashes (lines 1245-1251):

```python

# Inside incremental_update (lines 1245-1251)

file_hash = hashlib.sha256(file_path.read_bytes()).hexdigest()
stored_hash = store.get_file_hash(rel_path)
if file_hash == stored_hash:
    continue  # Skip: file unchanged despite VCS report

```

This catches cases where timestamps differ but content is identical — common after Git operations like rebases or branch switches.

### Parallel vs. Serial Parsing

The engine automatically selects execution strategy (lines 1260-1298):

- **Serial parsing** — when `CRG_SERIAL_PARSE=1` or file count is small
- **Parallel parsing** — using `_make_executor()` with kind selected by `_select_executor_kind()` (processes for CPU-bound work, threads for I/O-bound)

### Worker Function

Each file is processed by `_parse_single_file()` (lines 1028-1050):

```python
def _parse_single_file(repo_root: Path, rel_path: str, file_hash: str) -> ParsedResult:
    """
    Thread-local parser creation, file reading, and AST analysis.
    Returns nodes, edges, and computed hash for persistence.
    """
    full_path = repo_root / rel_path
    source_bytes = full_path.read_bytes()
    
    # Lazy thread-local parser initialization

    if not hasattr(_local, 'parser'):
        _local.parser = CodeParser()
    
    nodes, edges = _local.parser.parse(rel_path, source_bytes)
    return ParsedResult(nodes=nodes, edges=edges, file_hash=file_hash)

```

### Database Operations

Results are persisted through `store.store_file_nodes_edges()` (lines 1268-1285), with:

- **Single transaction commit** after all files complete
- **Permanent removal** of deleted files via `store.remove_files_permanently()` (lines 1299-1302)
- **Metadata refresh** for `branch`, `head_sha`, and `last_updated` timestamp (lines 1304-1307)

### Conditional Resolver Execution

Language-specific post-processing runs only when needed (lines 1310-1332):

```python

# Resolver gating example

if any(f.endswith('.py') for f in all_files):
    run_python_import_resolver(store)
if any('spring' in f.lower() for f in all_files):
    run_spring_di_resolver(store)

```

---

## Running Incremental Updates

### CLI Usage

```bash

# From repository root

crg update          # Uses stored HEAD SHA as diff base

# Force full rebuild

crg update --full

```

### Programmatic API

```python
from code_review_graph.incremental import incremental_update, get_changed_files
from code_review_graph.graph import GraphStore
from pathlib import Path

repo_root = Path(".")
store = GraphStore(repo_root / ".code-review-graph/graph.db")

# Basic incremental update

result = incremental_update(repo_root, store)
print(f"Updated {result['files_updated']} files")

# Manual change detection for custom logic

base = resolve_incremental_base(repo_root, store) or "HEAD~1"
changed = get_changed_files(repo_root, base)
print(f"VCS detected: {len(changed)} changed files")

```

### Expanding Impact Sets Manually

```python
from code_review_graph.incremental import find_dependents

changed = ["src/services/user.ts", "src/models/order.py"]

for rel_path in changed:
    dependents = find_dependents(store, rel_path)
    impact_set = {rel_path} | dependents
    print(f"{rel_path}: {len(dependents)} dependents, {len(impact_set)} total to process")

```

---

## Summary

- **Diff base resolution** (`resolve_incremental_base`, lines 644-667) validates stored SHAs and falls back to full rebuilds when history is unreachable
- **VCS integration** (`get_changed_files`, lines 670-708) handles Git and SVN with NUL-separated output parsing for rename detection
- **Impact expansion** (`find_dependents`, lines 986-1024) walks up to `_MAX_DEPENDENT_HOPS` (default 2) to find referencing files
- **Content hashing** prevents re-parsing of identical files even when VCS timestamps differ
- **Adaptive execution** chooses serial or parallel parsing based on workload size and environment variables
- **Conditional resolvers** avoid running language-specific analysis when no relevant files changed

The incremental update engine in code-review-graph achieves near-instant graph synchronization on large codebases by treating the graph itself as a cache and minimizing all operations — I/O, parsing, and database writes — to the theoretical minimum required for correctness.

---

## Frequently Asked Questions

### How does the incremental update engine handle Git renames?

The engine detects renames through `git diff --name-status -z`, which emits two paths for `R` (rename) and `C` (copy) statuses. The `_decode_name_status_paths()` function (lines 780-802) extracts both the old and new paths, ensuring the old file's nodes are removed from the graph while the new file is parsed fresh. This prevents orphaned nodes and broken references according to the tirth8205/code-review-graph source code.

### What triggers a full rebuild instead of an incremental update?

A full rebuild occurs when `resolve_incremental_base()` returns `None`, which happens if: (1) no `git_head_sha` is stored in the database, (2) the stored SHA no longer exists in the repository (e.g., after history rewriting), or (3) `_commit_object_exists()` fails to locate the commit object. The safety check against `_SAFE_GIT_REF` (lines 775-784) also rejects malformed references that could indicate corruption or tampering.

### Why does the engine hash file contents after already running Git diff?

Git's diff relies on metadata like timestamps and mode bits, which can change without content changes — common after `git rebase`, `git cherry-pick`, or branch switches. The SHA-256 content hash check (lines 1245-1251) provides a cryptographic guarantee of content equality, skipping parsing for files that are bit-for-bit identical to their stored versions. This second-layer deduplication significantly reduces parsing overhead in active development workflows.

### How does parallel parsing avoid race conditions?

The `_parse_single_file()` function (lines 1028-1050) uses **thread-local storage** via `_local` to create isolated `CodeParser` instances per thread. Database writes are batched and committed in a **single main-thread transaction** after all workers complete, preventing concurrent access to the SQLite store. The executor kind (process vs. thread) is selected by `_select_executor_kind()` based on whether the workload is CPU-bound or I/O-bound.