How Code-Review-Graph Handles Incremental Graph Updates by Re-Parsing Only Changed Files
Code-Review-Graph performs incremental graph updates by detecting changed files through Git diff, finding their dependents up to two hops, filtering unchanged files via SHA-256 hash comparison, and re-parsing only the necessary subset while leaving the rest of the graph untouched.
The code-review-graph repository by tirth8205 maintains a persistent SQLite-backed graph of code entities—functions, classes, imports, and dependencies. Rather than rebuilding this entire structure on every change, the tool implements a sophisticated incremental workflow in code_review_graph/incremental.py that minimizes re-parsing to only what has actually changed.
How the Incremental Update Process Works
The incremental engine follows an eight-step pipeline designed to surgically update the graph without touching stable portions of the codebase.
Step 1: Select a Diff Base
The process begins with resolve_incremental_base, which retrieves the last stored Git commit SHA (git_head_sha) from the GraphStore. The function validates that this commit still exists in the repository via _commit_object_exists. If the repository lacks Git metadata or the stored SHA is no longer valid, it automatically falls back to "HEAD~1":
# From code_review_graph/incremental.py:44-61
# Resolves the base commit for git diff comparison
def resolve_incremental_base(repo_root: Path, store: GraphStore) -> str:
stored_sha = store.get_metadata("git_head_sha")
if stored_sha and _commit_object_exists(repo_root, stored_sha):
return stored_sha
# Fallback when stored SHA is missing or invalid
return "HEAD~1"
Step 2: Detect Changed Files
The get_changed_files function executes git diff --name-status -z <base> -- to capture every file modification. The -z flag enables null-terminated output parsing, which correctly handles filenames containing spaces and captures both the old and new paths for renamed files:
# From code_review_graph/incremental.py:78-87
# Parses git diff output to identify changed, added, deleted, and renamed files
def get_changed_files(repo_root: Path, base: str) -> List[FileChange]:
cmd = ["git", "diff", "--name-status", "-z", base, "--"]
# ... decoding logic for status + old/new path pairs
Step 3: Reconcile Stale Files
Files that have disappeared from the repository or are now matched by ignore patterns are purged from the graph by _reconcile_stale_files. This prevents phantom nodes from accumulating across incremental updates:
# From code_review_graph/incremental.py:913-940
# Removes graph entries for files no longer present in the workspace
def _reconcile_stale_files(repo_root: Path, store: GraphStore) -> int:
# Scans filesystem vs. graph, deletes orphaned nodes/edges
...
Step 4: Find Dependent Files
For every changed file, find_dependents performs a breadth-first traversal up to two hops (configurable via the CRG_DEPENDENT_HOPS environment variable). This identifies all files that import from or otherwise reference the changed files—critical for capturing semantic impacts beyond direct modifications:
# From code_review_graph/incremental.py:885-910
# Traverses the dependency graph to find files affected by changes
def find_dependents(store: GraphStore, changed_file_ids: Set[str]) -> Set[str]:
hops = int(os.getenv("CRG_DEPENDENT_HOPS", "2"))
for _ in range(hops):
# Expand to importers of current frontier
...
Step 5: Build the Optimized Work-Set
The union of changed files + dependents undergoes aggressive filtering before parsing:
- Ignore patterns: Loaded via
_load_ignore_patternsand evaluated by_should_ignore - Binary detection: Skipped via
_is_binary - Content hashing: SHA-256 comparison in
code_review_graph/incremental.py:1232-1252skips files whose contents match the stored hash
# From code_review_graph/incremental.py:1232-1252
# Filters work-set through ignore patterns, binary check, and hash comparison
def _build_work_set(
repo_root: Path,
store: GraphStore,
candidate_files: Set[Path]
) -> List[Path]:
patterns = _load_ignore_patterns(repo_root)
for f in candidate_files:
if _should_ignore(f, patterns) or _is_binary(f):
continue
current_hash = hashlib.sha256(f.read_bytes()).hexdigest()
if current_hash == store.get_file_hash(f):
continue # Unchanged since last parse
work_set.append(f)
return work_set
Step 6: Parse Only Necessary Files
Parsing execution adapts to repository scale:
| Mode | Trigger | Implementation |
|---|---|---|
| Serial | Small repos or CRG_SERIAL_PARSE=1 |
Direct loop over work-set |
| Parallel | Default for larger repos | ProcessPoolExecutor or ThreadPoolExecutor selected by _select_executor_kind |
The module-level function _parse_single_file handles individual file parsing. Results feed into store.store_file_nodes_edges, which atomically writes nodes, edges, and the content hash for future incremental checks:
# From code_review_graph/incremental.py:330-340
# Parses files using Tree-Sitter, returns structured nodes and edges
def _parse_single_file(file_path: Path, language: str) -> ParseResult:
parser = CodeParser.for_language(language)
return parser.parse(file_path.read_text(), file_path)
Step 7: Update Metadata
Upon successful completion, the store records:
last_updated— timestamp of this updatelast_build_type="incremental"— distinguishes from full builds- Current VCS information via
_store_vcs_metadata
Step 8: Run Language-Specific Resolvers Conditionally
The final optimization: scoped resolvers execute only when relevant. The engine checks whether updated files belong to languages with registered resolvers (Python, ReScript, Spring, etc.) and invokes _run_python_resolver, _run_rescript_resolver, or equivalents only when needed:
# From code_review_graph/incremental.py:1288-1310
# Conditionally runs language-specific semantic analysis
def _maybe_run_resolvers(
store: GraphStore,
updated_files: List[Path]
) -> None:
languages = {detect_language(f) for f in updated_files}
if "python" in languages:
_run_python_resolver(store)
if "rescript" in languages:
_run_rescript_resolver(store)
# ... etc.
Practical Usage Examples
Programmatic API
from pathlib import Path
from code_review_graph.graph import GraphStore
from code_review_graph.incremental import incremental_update, full_build
repo_root = Path("/path/to/repo")
store = GraphStore(repo_root / ".code-review-graph/graph.db")
# Initial full build (run once per repository)
full_build(repo_root, store)
# Subsequent incremental updates after any change
result = incremental_update(
repo_root,
store,
base=None, # Auto-detect optimal diff base
changed_files=None, # Auto-detect via git diff
reconcile_stale=True, # Remove deleted/ignored files
)
print("Files re-parsed:", result["files_updated"])
print("New nodes added:", result["total_nodes"])
print("New edges added:", result["total_edges"])
Command-Line Interface
# Full build (one-time initialization)
$ crg build
# Incremental update (only changed files + dependents)
$ crg incremental
# Force specific diff base
$ crg incremental --base=HEAD~5
Key Source Files Supporting Incremental Updates
| File | Purpose |
|---|---|
code_review_graph/incremental.py |
Core incremental logic: change detection, dependency traversal, selective parsing |
code_review_graph/graph.py |
GraphStore SQLite implementation with node/edge queries and hash storage |
code_review_graph/parser.py |
CodeParser using Tree-Sitter for source-to-graph conversion |
code_review_graph/scoped_resolver.py |
Language-specific resolvers invoked conditionally per update |
Summary
- Diff-based change detection uses
git diff --name-statusto identify modified files with rename tracking - Dependency-aware expansion finds dependents up to
CRG_DEPENDENT_HOPS(default: 2) to capture semantic impacts - SHA-256 content hashing eliminates redundant parsing of unchanged files
- Parallel execution scales to large repositories via
ProcessPoolExecutororThreadPoolExecutor - Conditional resolver execution runs language-specific analysis only when relevant files change
- Persistent metadata stores commit SHAs and file hashes to enable future incremental comparisons
Frequently Asked Questions
How does code-review-graph know which files to skip during an incremental update?
The engine compares SHA-256 hashes of file contents against stored hashes in the GraphStore. Files whose hashes match their previous values are excluded from the work-set, even if Git reports them as touched (e.g., by a timestamp-only change). This check runs in code_review_graph/incremental.py lines 1232-1252 alongside ignore-pattern and binary-file filtering.
Can I control how many dependency hops the incremental parser follows?
Yes. Set the CRG_DEPENDENT_HOPS environment variable to configure how many levels of importers to include. The default value of 2 balances completeness with performance—sufficient to catch most breaking changes without exploding the work-set. Increasing this value captures deeper transitive dependencies at the cost of additional parsing.
What happens if code-review-graph hasn't seen a repository before?
The first run requires full_build() to establish the baseline graph and populate git_head_sha metadata. Subsequent runs use incremental_update(), which automatically detects the last processed commit and performs delta-based updates. If the stored commit SHA is missing or invalid, the system gracefully falls back to HEAD~1 as the diff base.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →