How to Handle Incremental Parsing for Large Codebases: A Deep Dive into code-graph-rag's Architecture

The code-graph-rag project solves incremental parsing by storing file hashes in the graph, re-parsing only changed files, and re-hydrating derived relations from persisted state—delivering full-index accuracy in seconds instead of minutes.

Large codebases present a fundamental challenge for code analysis systems: re-parsing every file on every run becomes prohibitively expensive as repositories grow to tens of thousands of files. The code-graph-rag repository implements a sophisticated incremental parsing system that maintains a complete, accurate graph without the cost of full re-indexing. This article examines the implementation details, from change detection to graph re-hydration.

Core Architecture of Incremental Parsing

The system centers on four coordinated components that work together to minimize work while guaranteeing correctness.

GraphUpdater: The Incremental Orchestrator

The GraphUpdater class in codebase_rag/graph_updater.py (line 228) manages the complete incremental workflow. It computes file changes, dispatches selective re-parsing, and restores derived structures that were stripped during the process.

from codebase_rag.graph_updater import GraphUpdater
from codebase_rag.storage import GraphStorage

store = GraphStorage("/tmp/cgr-store")
graph = store.load()

updater = GraphUpdater(store, incremental=True)
updater.update()  # Only changed files are re-parsed

File-Hash Based Change Detection

Each source file's content hash is stored as a node property in the graph. On every run, the updater computes fresh hashes and compares them against persisted values.

The _update_file method (line 760 in graph_updater.py) implements this logic:

  • Hash match: Skip parsing, reuse existing AST nodes
  • Hash mismatch: Queue file for full re-parse

This approach catches all content changes—including those invisible to file modification timestamps.

AST Cache for CPU Efficiency

The ASTCache class in codebase_rag/ast_cache.py persists abstract syntax trees to disk. Unchanged files load their AST directly from cache, eliminating the heavy parsing step entirely.

This design keeps memory usage bounded: only the working set of changed files requires fresh tree structures in RAM.

Graph Re-Hydration for Derived Relations

Certain graph relationships cannot survive incremental updates intact. Class inheritance edges, cross-file call edges, and similar derived structures are purged before selective parsing to prevent stale data.

After parsing completes, the system restores these relations from persisted state through dedicated re-hydration methods:

  • _rehydrate_class_inheritance_from_graph (line 1354)
  • _rehydrate_call_edges

This guarantees that the final graph contains complete, accurate relationships even though only a subset of files were re-parsed.

The Five-Stage Incremental Workflow

Understanding the precise sequence of operations reveals how the system maintains correctness while minimizing work.

Stage 1: Detect Changed Files

The updater walks the repository and computes a hash for each source file. These hashes are compared against file_hash properties stored in graph.nodes. The result is a precise changed_files set.

Stage 2: Selective Re-Parsing

Only files in changed_files pass through language-specific parsers:

  • LibClang for C/C++
  • tree-sitter for Rust
  • Python AST module for Python
  • Additional parsers in codebase_rag/parsers/

Unchanged files bypass this stage entirely.

Stage 3: Cache Reuse

For each unchanged file, the persisted AST loads from ASTCache. The parsing step is eliminated; existing node structures attach directly to the working graph.

Stage 4: Graph Re-Hydration

Derived relations are restored through systematic graph traversal. The re-hydration methods read the complete persisted state and reconstruct edges that span changed and unchanged file boundaries.

Stage 5: Consistency Verification

The evals/incremental.py harness validates every incremental run. It executes a "neutral edit scenario"—typically a whitespace-only change—and asserts that the resulting graph matches a clean full-index run.

from evals.incremental import run_neutral_edit_scenario, compare_states

inc_graph, clean_graph = run_neutral_edit_scenario(repo_path)
assert inc_graph == clean_graph, "Incremental graph diverges from clean index"

This test suite prevented regression in issue #532 and runs across multiple languages (C++, Rust, Java, Python).

Performance Characteristics for Large Codebases

The incremental design delivers measurable advantages that compound as repositories scale.

Metric Full Re-Index Incremental Update Improvement
Typical runtime Minutes Seconds 10-100x
Memory allocation All file ASTs Changed files only Bounded by edit size
CPU utilization 100% parse workload Proportional to changes Minimal for small edits
Graph consistency Guaranteed Verified by test suite Equivalent correctness

The deterministic graph property is critical: after any sequence of incremental updates, the graph matches what a fresh full-index would produce. This eliminates "drift" that accumulates in systems with approximate incremental logic.

Command-Line and Programmatic Interfaces

The system exposes incremental parsing through both CLI and Python APIs.

CLI usage:

cgr update_repository --path /path/to/repo --incremental

Python programmatic interface:

from codebase_rag.graph_updater import GraphUpdater
from codebase_rag.storage import GraphStorage

store = GraphStorage("/tmp/cgr-store")
graph = store.load()

updater = GraphUpdater(store, incremental=True)
updater.update()

# Graph now reflects latest source state

Both interfaces accept the same incremental flag, ensuring consistent behavior across usage patterns.

Key Implementation Files

File Purpose
codebase_rag/graph_updater.py Core incremental engine with GraphUpdater class and re-hydration methods
codebase_rag/ast_cache.py Persistent AST storage for unchanged files
evals/incremental.py Correctness validation harness
codebase_rag/tests/test_incremental_*.py Edge case coverage for inheritance, cross-file calls, and polyglot scenarios
codebase_rag/parsers/* Language-specific parsing implementations

Summary

  • code-graph-rag implements incremental parsing through coordinated change detection, selective re-parsing, AST caching, and graph re-hydration.
  • The GraphUpdater class orchestrates the complete workflow in codebase_rag/graph_updater.py.
  • File-hash comparison in _update_file provides precise, content-based change detection.
  • ASTCache eliminates parsing overhead for unchanged files.
  • Re-hydration methods restore derived relations (_rehydrate_class_inheritance_from_graph, _rehydrate_call_edges) to maintain graph completeness.
  • The evals/incremental.py test suite guarantees that incremental results match clean full-index results.

Frequently Asked Questions

How does code-graph-rag detect which files changed?

The system stores a SHA-256 hash of each file's content as a node property in the graph. On every incremental run, it computes fresh hashes and compares them against stored values. This content-based approach catches changes that file timestamps miss, including modifications that preserve modification time and git operations that alter content.

What happens to cross-file relationships when only one file changes?

Cross-file edges (class inheritance, function calls, imports) are purged before incremental parsing and re-hydrated afterward. The _rehydrate_class_inheritance_from_graph and related methods traverse the persisted graph state to reconstruct these relationships from the complete file set. This ensures that changes in one file correctly propagate to edges involving unchanged files.

When should I use incremental mode versus full re-indexing?

Use incremental mode for routine development workflows where small changes accumulate over time. Perform a full re-index when the storage format changes, after manual graph editing, or when diagnostic tools suggest graph corruption. The evals/incremental.py harness can verify that incremental updates remain consistent with fresh indexes.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →