Incremental Updates vs. Full Rebuilds in Code Review Graphs: A Complete Guide
Incremental updates parse only changed files and their dependents, while full rebuilds re-parse the entire repository from scratch.
The code-review-graph project by tirth8205 builds a SQLite-backed graph representing symbols, imports, and calls in a codebase. Two distinct build strategies—full_build and incremental_update—determine how this graph is constructed and maintained.
When Each Strategy Runs
Full rebuilds execute when you explicitly need a fresh graph. They are typically invoked manually or scheduled for clean slate operations.
Incremental updates trigger automatically on VCS changes or when incremental_update() is called directly. They analyze the difference between the current state and a base commit (default: HEAD~1) to minimize work.
Scope of Parsing: Complete vs. Selective
The most fundamental difference lies in what gets parsed.
Full Rebuild Parses Everything
In code_review_graph/incremental.py, full_build calls collect_all_files at lines 66-68 to gather every parseable file:
from code_review_graph.incremental import full_build
from code_review_graph.graph import GraphStore
from pathlib import Path
repo = Path("/path/to/repo")
store = GraphStore(repo / ".code-review-graph/graph.db")
result = full_build(repo, store)
print(f"Parsed {result['files_parsed']} files")
No file is skipped. The function rebuilds the entire graph topology from ground zero.
Incremental Update Parses Changed Files Plus Dependents
The incremental_update function at lines 196-215 builds a targeted file set:
- Changed files from
get_changed_files(VCS diff) - Dependent files from
find_dependents(graph traversal)
from code_review_graph.incremental import incremental_update
result = incremental_update(repo, store)
print(f"Updated {result['files_updated']} files "
f"({len(result['changed_files'])} changed, "
f"{len(result['dependent_files'])} dependents)")
Dependency expansion walks the graph up to a configurable hop limit (default: 2) to find files that import or call changed symbols—implemented at lines 90-102.
Hash-Based Skipping in Incremental Updates
Before parsing any candidate file, incremental_update computes a SHA-256 hash and compares it to stored values. Matching hashes skip reprocessing entirely (lines 44-50):
# Conceptual flow—hash check happens internally
if stored_hash == compute_sha256(filepath):
continue # File unchanged, skip parsing
Full rebuilds never perform hash checks—every file is always parsed.
Stale File Handling
Both strategies can detect and remove stale entries, but with different defaults.
| Strategy | Stale Reconciliation | Behavior |
|---|---|---|
full_build |
Always runs | _reconcile_stale_files executes before parsing (lines 13-20) |
incremental_update |
Optional (reconcile_stale=True) |
Can skip for speed when files are only modified, not deleted |
Skip reconciliation explicitly when performance matters:
result = incremental_update(repo, store, reconcile_stale=False)
Resolver Execution: All vs. Targeted
Language-specific resolvers (Python, ReScript, Spring, etc.) process parsed symbols into resolved graph edges.
- Full rebuild: Runs all resolvers after parsing (lines 145-158)
- Incremental update: Runs resolvers only if files of that language changed (lines 108-124)
For example, if only .py files changed, only the Python resolver executes—dramatically reducing computation.
Version Compatibility and Fallback Behavior
incremental_update includes a safety check at lines 71-79. If the stored C++ identity version mismatches the current binary, it automatically falls back to full_build before continuing. This prevents corruption from schema changes.
Metadata distinguishes the outcomes:
- Full rebuild:
last_build_type = "full" - Incremental update:
last_build_type = "incremental"
Performance Comparison by Repository State
| Scenario | Recommended Strategy | Rationale |
|---|---|---|
| Fresh clone or schema change | full_build |
No existing graph to increment from |
| Active development (few files changed) | incremental_update |
Minimal parsing, hash caching |
| Large refactoring across modules | full_build |
Dependency expansion may approach full scope anyway |
| CI pipeline with known base commit | incremental_update(base="HEAD~10") |
Explicit diff control |
Specifying a Custom Diff Base
For CI/CD scenarios, pin the comparison commit:
result = incremental_update(repo, store, base="HEAD~3")
This passes directly to get_changed_files, which executes git diff against that ref (lines 70-84).
Key Source Files
Understanding these implementations requires examining:
code_review_graph/incremental.py– Core algorithms for both build pathscode_review_graph/graph.py–GraphStoreSQLite persistence layerREADME.md– Visual flow diagram of incremental processing
Summary
- Full rebuilds exhaustively parse every file and run every resolver, producing a clean graph from scratch
- Incremental updates leverage VCS diffs, hash caching, and dependency expansion to minimize work
- Hash-based skipping prevents re-parsing unchanged files
- Resolver execution is targeted by language in incremental mode
- Version mismatches automatically trigger full rebuild fallback
- The
baseandreconcile_staleparameters provide fine-grained control
Frequently Asked Questions
When should I use a full rebuild instead of an incremental update?
Use full_build when the graph schema changes, after cloning a fresh repository, or when you suspect incremental drift. According to the source code, incremental_update automatically falls back to full rebuild if the C++ identity version mismatches—otherwise, incremental updates are preferred for speed.
How does incremental update find files that need re-parsing?
It combines two sources: get_changed_files (VCS diff) and find_dependents (graph traversal). The dependent search walks up to 2 hops by default to locate files importing or calling changed symbols. This limit is configurable in the find_dependents implementation at lines 90-102.
Can incremental updates miss changes that full rebuilds would catch?
Only if dependency expansion hops are insufficient for your codebase's coupling depth. The default 2-hop limit captures most indirect dependencies. For deeply intertwined modules, increase the hop limit or periodically run full_build to validate graph integrity.
What happens if I delete files and run incremental update with reconcile_stale=False?
Stale entries persist in the graph database. The _reconcile_stale_files function (lines 13-20) normally removes these, but setting reconcile_stale=False skips this step for speed. Use this optimization only when you know files are modified, not deleted or renamed.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →