How to Handle Incremental Parsing for Large Codebases: A Deep Dive into code-graph-rag's Architecture
The code-graph-rag project solves incremental parsing by storing file hashes in the graph, re-parsing only changed files, and re-hydrating derived relations from persisted state—delivering full-index accuracy in seconds instead of minutes.
Large codebases present a fundamental challenge for code analysis systems: re-parsing every file on every run becomes prohibitively expensive as repositories grow to tens of thousands of files. The code-graph-rag repository implements a sophisticated incremental parsing system that maintains a complete, accurate graph without the cost of full re-indexing. This article examines the implementation details, from change detection to graph re-hydration.
Core Architecture of Incremental Parsing
The system centers on four coordinated components that work together to minimize work while guaranteeing correctness.
GraphUpdater: The Incremental Orchestrator
The GraphUpdater class in codebase_rag/graph_updater.py (line 228) manages the complete incremental workflow. It computes file changes, dispatches selective re-parsing, and restores derived structures that were stripped during the process.
from codebase_rag.graph_updater import GraphUpdater
from codebase_rag.storage import GraphStorage
store = GraphStorage("/tmp/cgr-store")
graph = store.load()
updater = GraphUpdater(store, incremental=True)
updater.update() # Only changed files are re-parsed
File-Hash Based Change Detection
Each source file's content hash is stored as a node property in the graph. On every run, the updater computes fresh hashes and compares them against persisted values.
The _update_file method (line 760 in graph_updater.py) implements this logic:
- Hash match: Skip parsing, reuse existing AST nodes
- Hash mismatch: Queue file for full re-parse
This approach catches all content changes—including those invisible to file modification timestamps.
AST Cache for CPU Efficiency
The ASTCache class in codebase_rag/ast_cache.py persists abstract syntax trees to disk. Unchanged files load their AST directly from cache, eliminating the heavy parsing step entirely.
This design keeps memory usage bounded: only the working set of changed files requires fresh tree structures in RAM.
Graph Re-Hydration for Derived Relations
Certain graph relationships cannot survive incremental updates intact. Class inheritance edges, cross-file call edges, and similar derived structures are purged before selective parsing to prevent stale data.
After parsing completes, the system restores these relations from persisted state through dedicated re-hydration methods:
_rehydrate_class_inheritance_from_graph(line 1354)_rehydrate_call_edges
This guarantees that the final graph contains complete, accurate relationships even though only a subset of files were re-parsed.
The Five-Stage Incremental Workflow
Understanding the precise sequence of operations reveals how the system maintains correctness while minimizing work.
Stage 1: Detect Changed Files
The updater walks the repository and computes a hash for each source file. These hashes are compared against file_hash properties stored in graph.nodes. The result is a precise changed_files set.
Stage 2: Selective Re-Parsing
Only files in changed_files pass through language-specific parsers:
- LibClang for C/C++
- tree-sitter for Rust
- Python AST module for Python
- Additional parsers in
codebase_rag/parsers/
Unchanged files bypass this stage entirely.
Stage 3: Cache Reuse
For each unchanged file, the persisted AST loads from ASTCache. The parsing step is eliminated; existing node structures attach directly to the working graph.
Stage 4: Graph Re-Hydration
Derived relations are restored through systematic graph traversal. The re-hydration methods read the complete persisted state and reconstruct edges that span changed and unchanged file boundaries.
Stage 5: Consistency Verification
The evals/incremental.py harness validates every incremental run. It executes a "neutral edit scenario"—typically a whitespace-only change—and asserts that the resulting graph matches a clean full-index run.
from evals.incremental import run_neutral_edit_scenario, compare_states
inc_graph, clean_graph = run_neutral_edit_scenario(repo_path)
assert inc_graph == clean_graph, "Incremental graph diverges from clean index"
This test suite prevented regression in issue #532 and runs across multiple languages (C++, Rust, Java, Python).
Performance Characteristics for Large Codebases
The incremental design delivers measurable advantages that compound as repositories scale.
| Metric | Full Re-Index | Incremental Update | Improvement |
|---|---|---|---|
| Typical runtime | Minutes | Seconds | 10-100x |
| Memory allocation | All file ASTs | Changed files only | Bounded by edit size |
| CPU utilization | 100% parse workload | Proportional to changes | Minimal for small edits |
| Graph consistency | Guaranteed | Verified by test suite | Equivalent correctness |
The deterministic graph property is critical: after any sequence of incremental updates, the graph matches what a fresh full-index would produce. This eliminates "drift" that accumulates in systems with approximate incremental logic.
Command-Line and Programmatic Interfaces
The system exposes incremental parsing through both CLI and Python APIs.
CLI usage:
cgr update_repository --path /path/to/repo --incremental
Python programmatic interface:
from codebase_rag.graph_updater import GraphUpdater
from codebase_rag.storage import GraphStorage
store = GraphStorage("/tmp/cgr-store")
graph = store.load()
updater = GraphUpdater(store, incremental=True)
updater.update()
# Graph now reflects latest source state
Both interfaces accept the same incremental flag, ensuring consistent behavior across usage patterns.
Key Implementation Files
| File | Purpose |
|---|---|
codebase_rag/graph_updater.py |
Core incremental engine with GraphUpdater class and re-hydration methods |
codebase_rag/ast_cache.py |
Persistent AST storage for unchanged files |
evals/incremental.py |
Correctness validation harness |
codebase_rag/tests/test_incremental_*.py |
Edge case coverage for inheritance, cross-file calls, and polyglot scenarios |
codebase_rag/parsers/* |
Language-specific parsing implementations |
Summary
- code-graph-rag implements incremental parsing through coordinated change detection, selective re-parsing, AST caching, and graph re-hydration.
- The
GraphUpdaterclass orchestrates the complete workflow incodebase_rag/graph_updater.py. - File-hash comparison in
_update_fileprovides precise, content-based change detection. ASTCacheeliminates parsing overhead for unchanged files.- Re-hydration methods restore derived relations (
_rehydrate_class_inheritance_from_graph,_rehydrate_call_edges) to maintain graph completeness. - The
evals/incremental.pytest suite guarantees that incremental results match clean full-index results.
Frequently Asked Questions
How does code-graph-rag detect which files changed?
The system stores a SHA-256 hash of each file's content as a node property in the graph. On every incremental run, it computes fresh hashes and compares them against stored values. This content-based approach catches changes that file timestamps miss, including modifications that preserve modification time and git operations that alter content.
What happens to cross-file relationships when only one file changes?
Cross-file edges (class inheritance, function calls, imports) are purged before incremental parsing and re-hydrated afterward. The _rehydrate_class_inheritance_from_graph and related methods traverse the persisted graph state to reconstruct these relationships from the complete file set. This ensures that changes in one file correctly propagate to edges involving unchanged files.
When should I use incremental mode versus full re-indexing?
Use incremental mode for routine development workflows where small changes accumulate over time. Perform a full re-index when the storage format changes, after manual graph editing, or when diagnostic tools suggest graph corruption. The evals/incremental.py harness can verify that incremental updates remain consistent with fresh indexes.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →