Incremental Updates vs. Full Rebuilds in Code Review Graphs: A Complete Guide

Incremental updates parse only changed files and their dependents, while full rebuilds re-parse the entire repository from scratch.

The code-review-graph project by tirth8205 builds a SQLite-backed graph representing symbols, imports, and calls in a codebase. Two distinct build strategies—full_build and incremental_update—determine how this graph is constructed and maintained.

When Each Strategy Runs

Full rebuilds execute when you explicitly need a fresh graph. They are typically invoked manually or scheduled for clean slate operations.

Incremental updates trigger automatically on VCS changes or when incremental_update() is called directly. They analyze the difference between the current state and a base commit (default: HEAD~1) to minimize work.

Scope of Parsing: Complete vs. Selective

The most fundamental difference lies in what gets parsed.

Full Rebuild Parses Everything

In code_review_graph/incremental.py, full_build calls collect_all_files at lines 66-68 to gather every parseable file:

from code_review_graph.incremental import full_build
from code_review_graph.graph import GraphStore
from pathlib import Path

repo = Path("/path/to/repo")
store = GraphStore(repo / ".code-review-graph/graph.db")

result = full_build(repo, store)
print(f"Parsed {result['files_parsed']} files")

No file is skipped. The function rebuilds the entire graph topology from ground zero.

Incremental Update Parses Changed Files Plus Dependents

The incremental_update function at lines 196-215 builds a targeted file set:

  1. Changed files from get_changed_files (VCS diff)
  2. Dependent files from find_dependents (graph traversal)
from code_review_graph.incremental import incremental_update

result = incremental_update(repo, store)
print(f"Updated {result['files_updated']} files "
      f"({len(result['changed_files'])} changed, "
      f"{len(result['dependent_files'])} dependents)")

Dependency expansion walks the graph up to a configurable hop limit (default: 2) to find files that import or call changed symbols—implemented at lines 90-102.

Hash-Based Skipping in Incremental Updates

Before parsing any candidate file, incremental_update computes a SHA-256 hash and compares it to stored values. Matching hashes skip reprocessing entirely (lines 44-50):


# Conceptual flow—hash check happens internally

if stored_hash == compute_sha256(filepath):
    continue  # File unchanged, skip parsing

Full rebuilds never perform hash checks—every file is always parsed.

Stale File Handling

Both strategies can detect and remove stale entries, but with different defaults.

Strategy Stale Reconciliation Behavior
full_build Always runs _reconcile_stale_files executes before parsing (lines 13-20)
incremental_update Optional (reconcile_stale=True) Can skip for speed when files are only modified, not deleted

Skip reconciliation explicitly when performance matters:

result = incremental_update(repo, store, reconcile_stale=False)

Resolver Execution: All vs. Targeted

Language-specific resolvers (Python, ReScript, Spring, etc.) process parsed symbols into resolved graph edges.

  • Full rebuild: Runs all resolvers after parsing (lines 145-158)
  • Incremental update: Runs resolvers only if files of that language changed (lines 108-124)

For example, if only .py files changed, only the Python resolver executes—dramatically reducing computation.

Version Compatibility and Fallback Behavior

incremental_update includes a safety check at lines 71-79. If the stored C++ identity version mismatches the current binary, it automatically falls back to full_build before continuing. This prevents corruption from schema changes.

Metadata distinguishes the outcomes:

  • Full rebuild: last_build_type = "full"
  • Incremental update: last_build_type = "incremental"

Performance Comparison by Repository State

Scenario Recommended Strategy Rationale
Fresh clone or schema change full_build No existing graph to increment from
Active development (few files changed) incremental_update Minimal parsing, hash caching
Large refactoring across modules full_build Dependency expansion may approach full scope anyway
CI pipeline with known base commit incremental_update(base="HEAD~10") Explicit diff control

Specifying a Custom Diff Base

For CI/CD scenarios, pin the comparison commit:

result = incremental_update(repo, store, base="HEAD~3")

This passes directly to get_changed_files, which executes git diff against that ref (lines 70-84).

Key Source Files

Understanding these implementations requires examining:

Summary

  • Full rebuilds exhaustively parse every file and run every resolver, producing a clean graph from scratch
  • Incremental updates leverage VCS diffs, hash caching, and dependency expansion to minimize work
  • Hash-based skipping prevents re-parsing unchanged files
  • Resolver execution is targeted by language in incremental mode
  • Version mismatches automatically trigger full rebuild fallback
  • The base and reconcile_stale parameters provide fine-grained control

Frequently Asked Questions

When should I use a full rebuild instead of an incremental update?

Use full_build when the graph schema changes, after cloning a fresh repository, or when you suspect incremental drift. According to the source code, incremental_update automatically falls back to full rebuild if the C++ identity version mismatches—otherwise, incremental updates are preferred for speed.

How does incremental update find files that need re-parsing?

It combines two sources: get_changed_files (VCS diff) and find_dependents (graph traversal). The dependent search walks up to 2 hops by default to locate files importing or calling changed symbols. This limit is configurable in the find_dependents implementation at lines 90-102.

Can incremental updates miss changes that full rebuilds would catch?

Only if dependency expansion hops are insufficient for your codebase's coupling depth. The default 2-hop limit captures most indirect dependencies. For deeply intertwined modules, increase the hop limit or periodically run full_build to validate graph integrity.

What happens if I delete files and run incremental update with reconcile_stale=False?

Stale entries persist in the graph database. The _reconcile_stale_files function (lines 13-20) normally removes these, but setting reconcile_stale=False skips this step for speed. Use this optimization only when you know files are modified, not deleted or renamed.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →