How Code-Review-Graph Handles Incremental Graph Updates by Re-Parsing Only Changed Files

Code-Review-Graph performs incremental graph updates by detecting changed files through Git diff, finding their dependents up to two hops, filtering unchanged files via SHA-256 hash comparison, and re-parsing only the necessary subset while leaving the rest of the graph untouched.

The code-review-graph repository by tirth8205 maintains a persistent SQLite-backed graph of code entities—functions, classes, imports, and dependencies. Rather than rebuilding this entire structure on every change, the tool implements a sophisticated incremental workflow in code_review_graph/incremental.py that minimizes re-parsing to only what has actually changed.

How the Incremental Update Process Works

The incremental engine follows an eight-step pipeline designed to surgically update the graph without touching stable portions of the codebase.

Step 1: Select a Diff Base

The process begins with resolve_incremental_base, which retrieves the last stored Git commit SHA (git_head_sha) from the GraphStore. The function validates that this commit still exists in the repository via _commit_object_exists. If the repository lacks Git metadata or the stored SHA is no longer valid, it automatically falls back to "HEAD~1":


# From code_review_graph/incremental.py:44-61

# Resolves the base commit for git diff comparison

def resolve_incremental_base(repo_root: Path, store: GraphStore) -> str:
    stored_sha = store.get_metadata("git_head_sha")
    if stored_sha and _commit_object_exists(repo_root, stored_sha):
        return stored_sha
    # Fallback when stored SHA is missing or invalid

    return "HEAD~1"

Step 2: Detect Changed Files

The get_changed_files function executes git diff --name-status -z <base> -- to capture every file modification. The -z flag enables null-terminated output parsing, which correctly handles filenames containing spaces and captures both the old and new paths for renamed files:


# From code_review_graph/incremental.py:78-87

# Parses git diff output to identify changed, added, deleted, and renamed files

def get_changed_files(repo_root: Path, base: str) -> List[FileChange]:
    cmd = ["git", "diff", "--name-status", "-z", base, "--"]
    # ... decoding logic for status + old/new path pairs

Step 3: Reconcile Stale Files

Files that have disappeared from the repository or are now matched by ignore patterns are purged from the graph by _reconcile_stale_files. This prevents phantom nodes from accumulating across incremental updates:


# From code_review_graph/incremental.py:913-940

# Removes graph entries for files no longer present in the workspace

def _reconcile_stale_files(repo_root: Path, store: GraphStore) -> int:
    # Scans filesystem vs. graph, deletes orphaned nodes/edges

    ...

Step 4: Find Dependent Files

For every changed file, find_dependents performs a breadth-first traversal up to two hops (configurable via the CRG_DEPENDENT_HOPS environment variable). This identifies all files that import from or otherwise reference the changed files—critical for capturing semantic impacts beyond direct modifications:


# From code_review_graph/incremental.py:885-910

# Traverses the dependency graph to find files affected by changes

def find_dependents(store: GraphStore, changed_file_ids: Set[str]) -> Set[str]:
    hops = int(os.getenv("CRG_DEPENDENT_HOPS", "2"))
    for _ in range(hops):
        # Expand to importers of current frontier

        ...

Step 5: Build the Optimized Work-Set

The union of changed files + dependents undergoes aggressive filtering before parsing:

  • Ignore patterns: Loaded via _load_ignore_patterns and evaluated by _should_ignore
  • Binary detection: Skipped via _is_binary
  • Content hashing: SHA-256 comparison in code_review_graph/incremental.py:1232-1252 skips files whose contents match the stored hash

# From code_review_graph/incremental.py:1232-1252

# Filters work-set through ignore patterns, binary check, and hash comparison

def _build_work_set(
    repo_root: Path,
    store: GraphStore,
    candidate_files: Set[Path]
) -> List[Path]:
    patterns = _load_ignore_patterns(repo_root)
    for f in candidate_files:
        if _should_ignore(f, patterns) or _is_binary(f):
            continue
        current_hash = hashlib.sha256(f.read_bytes()).hexdigest()
        if current_hash == store.get_file_hash(f):
            continue  # Unchanged since last parse

        work_set.append(f)
    return work_set

Step 6: Parse Only Necessary Files

Parsing execution adapts to repository scale:

Mode Trigger Implementation
Serial Small repos or CRG_SERIAL_PARSE=1 Direct loop over work-set
Parallel Default for larger repos ProcessPoolExecutor or ThreadPoolExecutor selected by _select_executor_kind

The module-level function _parse_single_file handles individual file parsing. Results feed into store.store_file_nodes_edges, which atomically writes nodes, edges, and the content hash for future incremental checks:


# From code_review_graph/incremental.py:330-340

# Parses files using Tree-Sitter, returns structured nodes and edges

def _parse_single_file(file_path: Path, language: str) -> ParseResult:
    parser = CodeParser.for_language(language)
    return parser.parse(file_path.read_text(), file_path)

Step 7: Update Metadata

Upon successful completion, the store records:

  • last_updated — timestamp of this update
  • last_build_type="incremental" — distinguishes from full builds
  • Current VCS information via _store_vcs_metadata

Step 8: Run Language-Specific Resolvers Conditionally

The final optimization: scoped resolvers execute only when relevant. The engine checks whether updated files belong to languages with registered resolvers (Python, ReScript, Spring, etc.) and invokes _run_python_resolver, _run_rescript_resolver, or equivalents only when needed:


# From code_review_graph/incremental.py:1288-1310

# Conditionally runs language-specific semantic analysis

def _maybe_run_resolvers(
    store: GraphStore,
    updated_files: List[Path]
) -> None:
    languages = {detect_language(f) for f in updated_files}
    if "python" in languages:
        _run_python_resolver(store)
    if "rescript" in languages:
        _run_rescript_resolver(store)
    # ... etc.

Practical Usage Examples

Programmatic API

from pathlib import Path
from code_review_graph.graph import GraphStore
from code_review_graph.incremental import incremental_update, full_build

repo_root = Path("/path/to/repo")
store = GraphStore(repo_root / ".code-review-graph/graph.db")

# Initial full build (run once per repository)

full_build(repo_root, store)

# Subsequent incremental updates after any change

result = incremental_update(
    repo_root,
    store,
    base=None,              # Auto-detect optimal diff base

    changed_files=None,     # Auto-detect via git diff

    reconcile_stale=True,   # Remove deleted/ignored files

)

print("Files re-parsed:", result["files_updated"])
print("New nodes added:", result["total_nodes"])
print("New edges added:", result["total_edges"])

Command-Line Interface


# Full build (one-time initialization)

$ crg build

# Incremental update (only changed files + dependents)

$ crg incremental

# Force specific diff base

$ crg incremental --base=HEAD~5

Key Source Files Supporting Incremental Updates

File Purpose
code_review_graph/incremental.py Core incremental logic: change detection, dependency traversal, selective parsing
code_review_graph/graph.py GraphStore SQLite implementation with node/edge queries and hash storage
code_review_graph/parser.py CodeParser using Tree-Sitter for source-to-graph conversion
code_review_graph/scoped_resolver.py Language-specific resolvers invoked conditionally per update

Summary

  • Diff-based change detection uses git diff --name-status to identify modified files with rename tracking
  • Dependency-aware expansion finds dependents up to CRG_DEPENDENT_HOPS (default: 2) to capture semantic impacts
  • SHA-256 content hashing eliminates redundant parsing of unchanged files
  • Parallel execution scales to large repositories via ProcessPoolExecutor or ThreadPoolExecutor
  • Conditional resolver execution runs language-specific analysis only when relevant files change
  • Persistent metadata stores commit SHAs and file hashes to enable future incremental comparisons

Frequently Asked Questions

How does code-review-graph know which files to skip during an incremental update?

The engine compares SHA-256 hashes of file contents against stored hashes in the GraphStore. Files whose hashes match their previous values are excluded from the work-set, even if Git reports them as touched (e.g., by a timestamp-only change). This check runs in code_review_graph/incremental.py lines 1232-1252 alongside ignore-pattern and binary-file filtering.

Can I control how many dependency hops the incremental parser follows?

Yes. Set the CRG_DEPENDENT_HOPS environment variable to configure how many levels of importers to include. The default value of 2 balances completeness with performance—sufficient to catch most breaking changes without exploding the work-set. Increasing this value captures deeper transitive dependencies at the cost of additional parsing.

What happens if code-review-graph hasn't seen a repository before?

The first run requires full_build() to establish the baseline graph and populate git_head_sha metadata. Subsequent runs use incremental_update(), which automatically detects the last processed commit and performs delta-based updates. If the stored commit SHA is missing or invalid, the system gracefully falls back to HEAD~1 as the diff base.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →