# Incremental Updates vs. Full Rebuilds in Code Review Graphs: A Complete Guide

> Understand incremental updates vs full rebuilds in code review graphs. Learn how to optimize parsing for changed files and dependencies, boosting your code review efficiency.

- Repository: [Tirth Kanani/code-review-graph](https://github.com/tirth8205/code-review-graph)
- Tags: deep-dive
- Published: 2026-08-16

---

**Incremental updates parse only changed files and their dependents, while full rebuilds re-parse the entire repository from scratch.**

The `code-review-graph` project by tirth8205 builds a SQLite-backed graph representing symbols, imports, and calls in a codebase. Two distinct build strategies—`full_build` and `incremental_update`—determine how this graph is constructed and maintained.

## When Each Strategy Runs

**Full rebuilds** execute when you explicitly need a fresh graph. They are typically invoked manually or scheduled for clean slate operations.

**Incremental updates** trigger automatically on VCS changes or when `incremental_update()` is called directly. They analyze the difference between the current state and a base commit (default: `HEAD~1`) to minimize work.

## Scope of Parsing: Complete vs. Selective

The most fundamental difference lies in what gets parsed.

### Full Rebuild Parses Everything

In [`code_review_graph/incremental.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/incremental.py), `full_build` calls `collect_all_files` at lines 66-68 to gather every parseable file:

```python
from code_review_graph.incremental import full_build
from code_review_graph.graph import GraphStore
from pathlib import Path

repo = Path("/path/to/repo")
store = GraphStore(repo / ".code-review-graph/graph.db")

result = full_build(repo, store)
print(f"Parsed {result['files_parsed']} files")

```

No file is skipped. The function rebuilds the entire graph topology from ground zero.

### Incremental Update Parses Changed Files Plus Dependents

The `incremental_update` function at lines 196-215 builds a targeted file set:

1. **Changed files** from `get_changed_files` (VCS diff)
2. **Dependent files** from `find_dependents` (graph traversal)

```python
from code_review_graph.incremental import incremental_update

result = incremental_update(repo, store)
print(f"Updated {result['files_updated']} files "
      f"({len(result['changed_files'])} changed, "
      f"{len(result['dependent_files'])} dependents)")

```

Dependency expansion walks the graph up to a configurable hop limit (default: 2) to find files that import or call changed symbols—implemented at lines 90-102.

## Hash-Based Skipping in Incremental Updates

Before parsing any candidate file, `incremental_update` computes a SHA-256 hash and compares it to stored values. Matching hashes skip reprocessing entirely (lines 44-50):

```python

# Conceptual flow—hash check happens internally

if stored_hash == compute_sha256(filepath):
    continue  # File unchanged, skip parsing

```

Full rebuilds never perform hash checks—every file is always parsed.

## Stale File Handling

Both strategies can detect and remove stale entries, but with different defaults.

| Strategy | Stale Reconciliation | Behavior |
|----------|---------------------|----------|
| `full_build` | Always runs | `_reconcile_stale_files` executes before parsing (lines 13-20) |
| `incremental_update` | Optional (`reconcile_stale=True`) | Can skip for speed when files are only modified, not deleted |

Skip reconciliation explicitly when performance matters:

```python
result = incremental_update(repo, store, reconcile_stale=False)

```

## Resolver Execution: All vs. Targeted

Language-specific resolvers (Python, ReScript, Spring, etc.) process parsed symbols into resolved graph edges.

- **Full rebuild**: Runs *all* resolvers after parsing (lines 145-158)
- **Incremental update**: Runs resolvers *only* if files of that language changed (lines 108-124)

For example, if only `.py` files changed, only the Python resolver executes—dramatically reducing computation.

## Version Compatibility and Fallback Behavior

`incremental_update` includes a safety check at lines 71-79. If the stored C++ identity version mismatches the current binary, it automatically falls back to `full_build` before continuing. This prevents corruption from schema changes.

Metadata distinguishes the outcomes:

- Full rebuild: `last_build_type = "full"`
- Incremental update: `last_build_type = "incremental"`

## Performance Comparison by Repository State

| Scenario | Recommended Strategy | Rationale |
|----------|---------------------|-----------|
| Fresh clone or schema change | `full_build` | No existing graph to increment from |
| Active development (few files changed) | `incremental_update` | Minimal parsing, hash caching |
| Large refactoring across modules | `full_build` | Dependency expansion may approach full scope anyway |
| CI pipeline with known base commit | `incremental_update(base="HEAD~10")` | Explicit diff control |

## Specifying a Custom Diff Base

For CI/CD scenarios, pin the comparison commit:

```python
result = incremental_update(repo, store, base="HEAD~3")

```

This passes directly to `get_changed_files`, which executes `git diff` against that ref (lines 70-84).

## Key Source Files

Understanding these implementations requires examining:

- **[`code_review_graph/incremental.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/incremental.py)** – Core algorithms for both build paths
- **[`code_review_graph/graph.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/graph.py)** – `GraphStore` SQLite persistence layer
- **[`README.md`](https://github.com/tirth8205/code-review-graph/blob/main/README.md)** – Visual flow diagram of incremental processing

## Summary

- **Full rebuilds** exhaustively parse every file and run every resolver, producing a clean graph from scratch
- **Incremental updates** leverage VCS diffs, hash caching, and dependency expansion to minimize work
- Hash-based skipping prevents re-parsing unchanged files
- Resolver execution is targeted by language in incremental mode
- Version mismatches automatically trigger full rebuild fallback
- The `base` and `reconcile_stale` parameters provide fine-grained control

## Frequently Asked Questions

### When should I use a full rebuild instead of an incremental update?

Use `full_build` when the graph schema changes, after cloning a fresh repository, or when you suspect incremental drift. According to the source code, `incremental_update` automatically falls back to full rebuild if the C++ identity version mismatches—otherwise, incremental updates are preferred for speed.

### How does incremental update find files that need re-parsing?

It combines two sources: `get_changed_files` (VCS diff) and `find_dependents` (graph traversal). The dependent search walks up to 2 hops by default to locate files importing or calling changed symbols. This limit is configurable in the `find_dependents` implementation at lines 90-102.

### Can incremental updates miss changes that full rebuilds would catch?

Only if dependency expansion hops are insufficient for your codebase's coupling depth. The default 2-hop limit captures most indirect dependencies. For deeply intertwined modules, increase the hop limit or periodically run `full_build` to validate graph integrity.

### What happens if I delete files and run incremental update with `reconcile_stale=False`?

Stale entries persist in the graph database. The `_reconcile_stale_files` function (lines 13-20) normally removes these, but setting `reconcile_stale=False` skips this step for speed. Use this optimization only when you know files are modified, not deleted or renamed.