# How code-review-graph Supports Incremental Updates: A Technical Deep Dive

> Learn how code-review-graph supports incremental updates by efficiently detecting changes, finding dependents, and re-parsing only affected files using Git diff and SHA-256 hashing.

- Repository: [Tirth Kanani/code-review-graph](https://github.com/tirth8205/code-review-graph)
- Tags: deep-dive
- Published: 2026-08-13

---

**code-review-graph performs incremental updates by detecting changed files via Git diff, discovering their dependents through the existing graph, and re-parsing only the affected files while skipping unchanged content through SHA-256 hashing.**

code-review-graph maintains a knowledge graph representing symbols and relationships found in source repositories. Rebuilding the entire graph for every change would waste significant CPU and I/O resources, so the library implements a sophisticated **incremental update** system that minimizes work by targeting only changed files and their transitive dependencies. This ten-step process is implemented in [`code_review_graph/incremental.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/incremental.py) and leverages the existing graph structure to determine the minimal set of files requiring re-analysis.

## The Incremental Update Pipeline

The incremental workflow in [`code_review_graph/incremental.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/incremental.py) spans ten distinct steps, each designed to ensure accuracy while maximizing performance.

### Step 1: Repository Root Detection

The system locates the project root by traversing upward to find a `.git` or SVN marker. Users can override this behavior by setting the `CRG_REPO_ROOT` environment variable. The `find_repo_root()` and `find_project_root()` functions handle this discovery.

### Step 2: Diff Base Resolution

For Git repositories, the system uses `resolve_incremental_base()` to identify the last stored commit SHA (`git_head_sha`) as the comparison base. If this commit no longer exists in the history, the system falls back to a full rebuild to ensure consistency.

### Step 3: Changed File Discovery

The `get_changed_files()` function executes `git diff --name-status -z <base>` (or the SVN equivalent) to identify modified paths. The `_decode_name_status_paths()` helper ensures rename operations capture both old and new file paths.

### Step 4: Stale File Reconciliation

The `_reconcile_stale_files()` function removes nodes and edges for files that have been deleted or are now excluded by ignore patterns, preventing phantom relationships in the graph.

### Step 5: Dependent File Discovery

Using the existing graph, `find_dependents()` walks import and call edges to identify files depending on changed files. The `_single_hop_dependents()` helper performs bounded traversal limited by `_MAX_DEPENDENT_HOPS` (default 2) and capped at `_MAX_DEPENDENT_FILES` (500) to prevent analysis explosions.

### Step 6: Content-Based Filtering

Before parsing, the system compares SHA-256 hashes against stored values to skip files whose contents remain unchanged, even if they appear in the changed-file list due to renaming operations.

### Step 7: Parallel File Parsing

The `_parse_single_file()` function processes affected files using either a process pool or thread pool, selected by `_select_executor_kind()`. Workers reuse cached `CodeParser` instances to eliminate startup overhead. Process pools provide isolation by default, while thread pools prevent Windows deadlocks when using FastMCP stdio transport.

### Step 8: Graph Store Updates

New nodes and edges are persisted via `store_file_nodes_edges()`. Deletions are committed before insertions to avoid nested transaction errors in the SQLite-backed `GraphStore`.

### Step 9: Metadata Recording

The `_store_vcs_metadata()` function updates `last_updated`, `last_build_type`, branch name, and commit SHA, establishing the foundation for future incremental comparisons.

### Step 10: Language-Specific Resolution

Optional resolvers execute only for languages present in the changed set, including Python, ReScript, Spring, Temporal, and HCL modules.

## Performance Optimizations in Incremental Graph Building

The implementation includes several safeguards against pathological cases during **incremental updates**.

**Dependency Bounding:** The traversal limits prevent infinite walks through dense dependency chains by enforcing `_MAX_DEPENDENT_HOPS` and `_MAX_DEPENDENT_FILES` constraints.

**Hash-Based Deduplication:** SHA-256 content hashing eliminates redundant parsing of files unchanged despite appearing in Git's rename tracking.

**Transport-Aware Parallelism:** The executor selection logic automatically detects FastMCP stdio transport and switches from processes to threads on Windows to avoid deadlock conditions.

## Practical Implementation: Running Incremental Updates

The following examples demonstrate running incremental updates using the `code_review_graph` API.

Initialize a graph store and perform a full build first:

```python
from pathlib import Path
from code_review_graph.graph import GraphStore
from code_review_graph.incremental import incremental_update, full_build

# Initialize SQLite-backed GraphStore

store = GraphStore(Path("/tmp/myrepo/graph.db"))

# Full rebuild - typically run once after cloning

full_build(Path("/tmp/myrepo"), store)

```

After the initial build, perform **incremental updates** automatically:

```python

# Auto-detect changes and update incrementally

result = incremental_update(Path("/tmp/myrepo"), store)

print("Files re-parsed:", result["files_updated"])
print("New nodes added:", result["total_nodes"])
print("New edges added:", result["total_edges"])

```

You can also specify an explicit Git reference as the diff base:

```python
incremental_update(
    repo_root=Path("/tmp/myrepo"),
    store=store,
    base="v1.2.3",               # Any valid git ref

    changed_files=None,          # Compute diffs automatically

    reconcile_stale=True,
)

```

## Summary

- **Change Detection:** The system uses `git diff --name-status` via `get_changed_files()` to identify modifications against the last stored commit SHA.
- **Dependency Expansion:** `find_dependents()` performs bounded graph traversal to identify affected files without processing the entire repository.
- **Content Deduplication:** SHA-256 hashing prevents re-parsing unchanged files even when Git reports them as renamed.
- **Parallel Processing:** The `_select_executor_kind()` function chooses between process and thread pools based on the execution environment to maximize throughput.
- **Metadata Tracking:** VCS state is preserved through `_store_vcs_metadata()`, enabling reliable incremental bases across multiple update cycles.

## Frequently Asked Questions

### What happens if the base commit is missing from the repository?

If the stored `git_head_sha` no longer exists in the repository (for example, after a history rewrite or garbage collection), `resolve_incremental_base()` triggers a full rebuild. This ensures the graph maintains consistency with the actual repository state rather than attempting incremental updates from an unreachable reference.

### How does code-review-graph handle file renames during incremental updates?

The `_decode_name_status_paths()` function parses Git's rename detection output to capture both the old and new paths. Additionally, SHA-256 hash comparison identifies when file contents remain identical despite the rename, allowing the system to update path metadata without re-parsing the actual code.

### What limits prevent dependency explosion during incremental updates?

Two constants guard against excessive traversal: `_MAX_DEPENDENT_HOPS` defaults to 2 hops from changed files, and `_MAX_DEPENDENT_FILES` caps total dependent discovery at 500 files. These bounds ensure that changes to widely imported utility files do not trigger full-repository re-parsing.

### Can incremental updates run in parallel with other operations?

Yes. The parsing phase uses `_make_executor()` to create either process or thread pools based on the transport context. Process pools provide isolation for CPU-bound parsing, while thread pools prevent Windows-specific deadlocks when running under FastMCP stdio transport, allowing safe concurrent execution.