# How Code-Review-Graph Handles Incremental Graph Updates by Re-Parsing Only Changed Files

> Code-Review-Graph efficiently updates graphs by re-parsing only changed files. Learn how it uses Git diff and SHA-256 hashing for incremental updates.

- Repository: [Tirth Kanani/code-review-graph](https://github.com/tirth8205/code-review-graph)
- Tags: internals
- Published: 2026-08-14

---

**Code-Review-Graph performs incremental graph updates by detecting changed files through Git diff, finding their dependents up to two hops, filtering unchanged files via SHA-256 hash comparison, and re-parsing only the necessary subset while leaving the rest of the graph untouched.**

The **code-review-graph** repository by `tirth8205` maintains a persistent **SQLite-backed graph** of code entities—functions, classes, imports, and dependencies. Rather than rebuilding this entire structure on every change, the tool implements a sophisticated incremental workflow in [`code_review_graph/incremental.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/incremental.py) that minimizes re-parsing to only what has actually changed.

## How the Incremental Update Process Works

The incremental engine follows an eight-step pipeline designed to surgically update the graph without touching stable portions of the codebase.

### Step 1: Select a Diff Base

The process begins with `resolve_incremental_base`, which retrieves the last stored Git commit SHA (`git_head_sha`) from the `GraphStore`. The function validates that this commit still exists in the repository via `_commit_object_exists`. If the repository lacks Git metadata or the stored SHA is no longer valid, it automatically falls back to `"HEAD~1"`:

```python

# From code_review_graph/incremental.py:44-61

# Resolves the base commit for git diff comparison

def resolve_incremental_base(repo_root: Path, store: GraphStore) -> str:
    stored_sha = store.get_metadata("git_head_sha")
    if stored_sha and _commit_object_exists(repo_root, stored_sha):
        return stored_sha
    # Fallback when stored SHA is missing or invalid

    return "HEAD~1"

```

### Step 2: Detect Changed Files

The `get_changed_files` function executes `git diff --name-status -z <base> --` to capture every file modification. The `-z` flag enables null-terminated output parsing, which correctly handles filenames containing spaces and captures both the old and new paths for renamed files:

```python

# From code_review_graph/incremental.py:78-87

# Parses git diff output to identify changed, added, deleted, and renamed files

def get_changed_files(repo_root: Path, base: str) -> List[FileChange]:
    cmd = ["git", "diff", "--name-status", "-z", base, "--"]
    # ... decoding logic for status + old/new path pairs

```

### Step 3: Reconcile Stale Files

Files that have disappeared from the repository or are now matched by ignore patterns are purged from the graph by `_reconcile_stale_files`. This prevents phantom nodes from accumulating across incremental updates:

```python

# From code_review_graph/incremental.py:913-940

# Removes graph entries for files no longer present in the workspace

def _reconcile_stale_files(repo_root: Path, store: GraphStore) -> int:
    # Scans filesystem vs. graph, deletes orphaned nodes/edges

    ...

```

### Step 4: Find Dependent Files

For every changed file, `find_dependents` performs a **breadth-first traversal up to two hops** (configurable via the `CRG_DEPENDENT_HOPS` environment variable). This identifies all files that import from or otherwise reference the changed files—critical for capturing semantic impacts beyond direct modifications:

```python

# From code_review_graph/incremental.py:885-910

# Traverses the dependency graph to find files affected by changes

def find_dependents(store: GraphStore, changed_file_ids: Set[str]) -> Set[str]:
    hops = int(os.getenv("CRG_DEPENDENT_HOPS", "2"))
    for _ in range(hops):
        # Expand to importers of current frontier

        ...

```

### Step 5: Build the Optimized Work-Set

The union of **changed files + dependents** undergoes aggressive filtering before parsing:

- **Ignore patterns**: Loaded via `_load_ignore_patterns` and evaluated by `_should_ignore`
- **Binary detection**: Skipped via `_is_binary`
- **Content hashing**: SHA-256 comparison in `code_review_graph/incremental.py:1232-1252` skips files whose contents match the stored hash

```python

# From code_review_graph/incremental.py:1232-1252

# Filters work-set through ignore patterns, binary check, and hash comparison

def _build_work_set(
    repo_root: Path,
    store: GraphStore,
    candidate_files: Set[Path]
) -> List[Path]:
    patterns = _load_ignore_patterns(repo_root)
    for f in candidate_files:
        if _should_ignore(f, patterns) or _is_binary(f):
            continue
        current_hash = hashlib.sha256(f.read_bytes()).hexdigest()
        if current_hash == store.get_file_hash(f):
            continue  # Unchanged since last parse

        work_set.append(f)
    return work_set

```

### Step 6: Parse Only Necessary Files

Parsing execution adapts to repository scale:

| Mode | Trigger | Implementation |
|------|---------|----------------|
| **Serial** | Small repos or `CRG_SERIAL_PARSE=1` | Direct loop over work-set |
| **Parallel** | Default for larger repos | `ProcessPoolExecutor` or `ThreadPoolExecutor` selected by `_select_executor_kind` |

The module-level function `_parse_single_file` handles individual file parsing. Results feed into `store.store_file_nodes_edges`, which atomically writes nodes, edges, and the content hash for future incremental checks:

```python

# From code_review_graph/incremental.py:330-340

# Parses files using Tree-Sitter, returns structured nodes and edges

def _parse_single_file(file_path: Path, language: str) -> ParseResult:
    parser = CodeParser.for_language(language)
    return parser.parse(file_path.read_text(), file_path)

```

### Step 7: Update Metadata

Upon successful completion, the store records:

- `last_updated` — timestamp of this update
- `last_build_type="incremental"` — distinguishes from full builds
- Current VCS information via `_store_vcs_metadata`

### Step 8: Run Language-Specific Resolvers Conditionally

The final optimization: **scoped resolvers execute only when relevant**. The engine checks whether updated files belong to languages with registered resolvers (Python, ReScript, Spring, etc.) and invokes `_run_python_resolver`, `_run_rescript_resolver`, or equivalents only when needed:

```python

# From code_review_graph/incremental.py:1288-1310

# Conditionally runs language-specific semantic analysis

def _maybe_run_resolvers(
    store: GraphStore,
    updated_files: List[Path]
) -> None:
    languages = {detect_language(f) for f in updated_files}
    if "python" in languages:
        _run_python_resolver(store)
    if "rescript" in languages:
        _run_rescript_resolver(store)
    # ... etc.

```

## Practical Usage Examples

### Programmatic API

```python
from pathlib import Path
from code_review_graph.graph import GraphStore
from code_review_graph.incremental import incremental_update, full_build

repo_root = Path("/path/to/repo")
store = GraphStore(repo_root / ".code-review-graph/graph.db")

# Initial full build (run once per repository)

full_build(repo_root, store)

# Subsequent incremental updates after any change

result = incremental_update(
    repo_root,
    store,
    base=None,              # Auto-detect optimal diff base

    changed_files=None,     # Auto-detect via git diff

    reconcile_stale=True,   # Remove deleted/ignored files

)

print("Files re-parsed:", result["files_updated"])
print("New nodes added:", result["total_nodes"])
print("New edges added:", result["total_edges"])

```

### Command-Line Interface

```bash

# Full build (one-time initialization)

$ crg build

# Incremental update (only changed files + dependents)

$ crg incremental

# Force specific diff base

$ crg incremental --base=HEAD~5

```

## Key Source Files Supporting Incremental Updates

| File | Purpose |
|------|---------|
| [`code_review_graph/incremental.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/incremental.py) | Core incremental logic: change detection, dependency traversal, selective parsing |
| [`code_review_graph/graph.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/graph.py) | `GraphStore` SQLite implementation with node/edge queries and hash storage |
| [`code_review_graph/parser.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/parser.py) | `CodeParser` using Tree-Sitter for source-to-graph conversion |
| [`code_review_graph/scoped_resolver.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/scoped_resolver.py) | Language-specific resolvers invoked conditionally per update |

## Summary

- **Diff-based change detection** uses `git diff --name-status` to identify modified files with rename tracking
- **Dependency-aware expansion** finds dependents up to `CRG_DEPENDENT_HOPS` (default: 2) to capture semantic impacts
- **SHA-256 content hashing** eliminates redundant parsing of unchanged files
- **Parallel execution** scales to large repositories via `ProcessPoolExecutor` or `ThreadPoolExecutor`
- **Conditional resolver execution** runs language-specific analysis only when relevant files change
- **Persistent metadata** stores commit SHAs and file hashes to enable future incremental comparisons

## Frequently Asked Questions

### How does code-review-graph know which files to skip during an incremental update?

The engine compares SHA-256 hashes of file contents against stored hashes in the `GraphStore`. Files whose hashes match their previous values are excluded from the work-set, even if Git reports them as touched (e.g., by a timestamp-only change). This check runs in [`code_review_graph/incremental.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/incremental.py) lines 1232-1252 alongside ignore-pattern and binary-file filtering.

### Can I control how many dependency hops the incremental parser follows?

Yes. Set the `CRG_DEPENDENT_HOPS` environment variable to configure how many levels of importers to include. The default value of `2` balances completeness with performance—sufficient to catch most breaking changes without exploding the work-set. Increasing this value captures deeper transitive dependencies at the cost of additional parsing.

### What happens if code-review-graph hasn't seen a repository before?

The first run requires `full_build()` to establish the baseline graph and populate `git_head_sha` metadata. Subsequent runs use `incremental_update()`, which automatically detects the last processed commit and performs delta-based updates. If the stored commit SHA is missing or invalid, the system gracefully falls back to `HEAD~1` as the diff base.