How code-review-graph Supports Incremental Updates: A Technical Deep Dive
code-review-graph performs incremental updates by detecting changed files via Git diff, discovering their dependents through the existing graph, and re-parsing only the affected files while skipping unchanged content through SHA-256 hashing.
code-review-graph maintains a knowledge graph representing symbols and relationships found in source repositories. Rebuilding the entire graph for every change would waste significant CPU and I/O resources, so the library implements a sophisticated incremental update system that minimizes work by targeting only changed files and their transitive dependencies. This ten-step process is implemented in code_review_graph/incremental.py and leverages the existing graph structure to determine the minimal set of files requiring re-analysis.
The Incremental Update Pipeline
The incremental workflow in code_review_graph/incremental.py spans ten distinct steps, each designed to ensure accuracy while maximizing performance.
Step 1: Repository Root Detection
The system locates the project root by traversing upward to find a .git or SVN marker. Users can override this behavior by setting the CRG_REPO_ROOT environment variable. The find_repo_root() and find_project_root() functions handle this discovery.
Step 2: Diff Base Resolution
For Git repositories, the system uses resolve_incremental_base() to identify the last stored commit SHA (git_head_sha) as the comparison base. If this commit no longer exists in the history, the system falls back to a full rebuild to ensure consistency.
Step 3: Changed File Discovery
The get_changed_files() function executes git diff --name-status -z <base> (or the SVN equivalent) to identify modified paths. The _decode_name_status_paths() helper ensures rename operations capture both old and new file paths.
Step 4: Stale File Reconciliation
The _reconcile_stale_files() function removes nodes and edges for files that have been deleted or are now excluded by ignore patterns, preventing phantom relationships in the graph.
Step 5: Dependent File Discovery
Using the existing graph, find_dependents() walks import and call edges to identify files depending on changed files. The _single_hop_dependents() helper performs bounded traversal limited by _MAX_DEPENDENT_HOPS (default 2) and capped at _MAX_DEPENDENT_FILES (500) to prevent analysis explosions.
Step 6: Content-Based Filtering
Before parsing, the system compares SHA-256 hashes against stored values to skip files whose contents remain unchanged, even if they appear in the changed-file list due to renaming operations.
Step 7: Parallel File Parsing
The _parse_single_file() function processes affected files using either a process pool or thread pool, selected by _select_executor_kind(). Workers reuse cached CodeParser instances to eliminate startup overhead. Process pools provide isolation by default, while thread pools prevent Windows deadlocks when using FastMCP stdio transport.
Step 8: Graph Store Updates
New nodes and edges are persisted via store_file_nodes_edges(). Deletions are committed before insertions to avoid nested transaction errors in the SQLite-backed GraphStore.
Step 9: Metadata Recording
The _store_vcs_metadata() function updates last_updated, last_build_type, branch name, and commit SHA, establishing the foundation for future incremental comparisons.
Step 10: Language-Specific Resolution
Optional resolvers execute only for languages present in the changed set, including Python, ReScript, Spring, Temporal, and HCL modules.
Performance Optimizations in Incremental Graph Building
The implementation includes several safeguards against pathological cases during incremental updates.
Dependency Bounding: The traversal limits prevent infinite walks through dense dependency chains by enforcing _MAX_DEPENDENT_HOPS and _MAX_DEPENDENT_FILES constraints.
Hash-Based Deduplication: SHA-256 content hashing eliminates redundant parsing of files unchanged despite appearing in Git's rename tracking.
Transport-Aware Parallelism: The executor selection logic automatically detects FastMCP stdio transport and switches from processes to threads on Windows to avoid deadlock conditions.
Practical Implementation: Running Incremental Updates
The following examples demonstrate running incremental updates using the code_review_graph API.
Initialize a graph store and perform a full build first:
from pathlib import Path
from code_review_graph.graph import GraphStore
from code_review_graph.incremental import incremental_update, full_build
# Initialize SQLite-backed GraphStore
store = GraphStore(Path("/tmp/myrepo/graph.db"))
# Full rebuild - typically run once after cloning
full_build(Path("/tmp/myrepo"), store)
After the initial build, perform incremental updates automatically:
# Auto-detect changes and update incrementally
result = incremental_update(Path("/tmp/myrepo"), store)
print("Files re-parsed:", result["files_updated"])
print("New nodes added:", result["total_nodes"])
print("New edges added:", result["total_edges"])
You can also specify an explicit Git reference as the diff base:
incremental_update(
repo_root=Path("/tmp/myrepo"),
store=store,
base="v1.2.3", # Any valid git ref
changed_files=None, # Compute diffs automatically
reconcile_stale=True,
)
Summary
- Change Detection: The system uses
git diff --name-statusviaget_changed_files()to identify modifications against the last stored commit SHA. - Dependency Expansion:
find_dependents()performs bounded graph traversal to identify affected files without processing the entire repository. - Content Deduplication: SHA-256 hashing prevents re-parsing unchanged files even when Git reports them as renamed.
- Parallel Processing: The
_select_executor_kind()function chooses between process and thread pools based on the execution environment to maximize throughput. - Metadata Tracking: VCS state is preserved through
_store_vcs_metadata(), enabling reliable incremental bases across multiple update cycles.
Frequently Asked Questions
What happens if the base commit is missing from the repository?
If the stored git_head_sha no longer exists in the repository (for example, after a history rewrite or garbage collection), resolve_incremental_base() triggers a full rebuild. This ensures the graph maintains consistency with the actual repository state rather than attempting incremental updates from an unreachable reference.
How does code-review-graph handle file renames during incremental updates?
The _decode_name_status_paths() function parses Git's rename detection output to capture both the old and new paths. Additionally, SHA-256 hash comparison identifies when file contents remain identical despite the rename, allowing the system to update path metadata without re-parsing the actual code.
What limits prevent dependency explosion during incremental updates?
Two constants guard against excessive traversal: _MAX_DEPENDENT_HOPS defaults to 2 hops from changed files, and _MAX_DEPENDENT_FILES caps total dependent discovery at 500 files. These bounds ensure that changes to widely imported utility files do not trigger full-repository re-parsing.
Can incremental updates run in parallel with other operations?
Yes. The parsing phase uses _make_executor() to create either process or thread pools based on the transport context. Process pools provide isolation for CPU-bound parsing, while thread pools prevent Windows-specific deadlocks when running under FastMCP stdio transport, allowing safe concurrent execution.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →