# How to Handle Incremental Parsing for Large Codebases: A Deep Dive into code-graph-rag's Architecture

> Discover how code-graph-rag handles incremental parsing for large codebases efficiently. Learn to re-parse only changed files and maintain full-index accuracy in seconds.

- Repository: [Vitali Avagyan/code-graph-rag](https://github.com/vitali87/code-graph-rag)
- Tags: deep-dive
- Published: 2026-08-20

---

**The code-graph-rag project solves incremental parsing by storing file hashes in the graph, re-parsing only changed files, and re-hydrating derived relations from persisted state—delivering full-index accuracy in seconds instead of minutes.**

Large codebases present a fundamental challenge for code analysis systems: re-parsing every file on every run becomes prohibitively expensive as repositories grow to tens of thousands of files. The **code-graph-rag** repository implements a sophisticated incremental parsing system that maintains a complete, accurate graph without the cost of full re-indexing. This article examines the implementation details, from change detection to graph re-hydration.

## Core Architecture of Incremental Parsing

The system centers on four coordinated components that work together to minimize work while guaranteeing correctness.

### GraphUpdater: The Incremental Orchestrator

The `GraphUpdater` class in [`codebase_rag/graph_updater.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/graph_updater.py) (line 228) manages the complete incremental workflow. It computes file changes, dispatches selective re-parsing, and restores derived structures that were stripped during the process.

```python
from codebase_rag.graph_updater import GraphUpdater
from codebase_rag.storage import GraphStorage

store = GraphStorage("/tmp/cgr-store")
graph = store.load()

updater = GraphUpdater(store, incremental=True)
updater.update()  # Only changed files are re-parsed

```

### File-Hash Based Change Detection

Each source file's **content hash** is stored as a node property in the graph. On every run, the updater computes fresh hashes and compares them against persisted values.

The `_update_file` method (line 760 in [`graph_updater.py`](https://github.com/vitali87/code-graph-rag/blob/main/graph_updater.py)) implements this logic:

- **Hash match**: Skip parsing, reuse existing AST nodes
- **Hash mismatch**: Queue file for full re-parse

This approach catches all content changes—including those invisible to file modification timestamps.

### AST Cache for CPU Efficiency

The `ASTCache` class in [`codebase_rag/ast_cache.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/ast_cache.py) persists abstract syntax trees to disk. Unchanged files load their AST directly from cache, eliminating the heavy parsing step entirely.

This design keeps memory usage bounded: only the working set of changed files requires fresh tree structures in RAM.

### Graph Re-Hydration for Derived Relations

Certain graph relationships cannot survive incremental updates intact. **Class inheritance edges**, **cross-file call edges**, and similar derived structures are purged before selective parsing to prevent stale data.

After parsing completes, the system restores these relations from persisted state through dedicated re-hydration methods:

- `_rehydrate_class_inheritance_from_graph` (line 1354)
- `_rehydrate_call_edges`

This guarantees that the final graph contains complete, accurate relationships even though only a subset of files were re-parsed.

## The Five-Stage Incremental Workflow

Understanding the precise sequence of operations reveals how the system maintains correctness while minimizing work.

### Stage 1: Detect Changed Files

The updater walks the repository and computes a hash for each source file. These hashes are compared against `file_hash` properties stored in `graph.nodes`. The result is a precise `changed_files` set.

### Stage 2: Selective Re-Parsing

Only files in `changed_files` pass through language-specific parsers:

- **LibClang** for C/C++
- **tree-sitter** for Rust
- **Python AST module** for Python
- Additional parsers in `codebase_rag/parsers/`

Unchanged files bypass this stage entirely.

### Stage 3: Cache Reuse

For each unchanged file, the persisted AST loads from `ASTCache`. The parsing step is eliminated; existing node structures attach directly to the working graph.

### Stage 4: Graph Re-Hydration

Derived relations are restored through systematic graph traversal. The re-hydration methods read the complete persisted state and reconstruct edges that span changed and unchanged file boundaries.

### Stage 5: Consistency Verification

The [`evals/incremental.py`](https://github.com/vitali87/code-graph-rag/blob/main/evals/incremental.py) harness validates every incremental run. It executes a "neutral edit scenario"—typically a whitespace-only change—and asserts that the resulting graph matches a clean full-index run.

```python
from evals.incremental import run_neutral_edit_scenario, compare_states

inc_graph, clean_graph = run_neutral_edit_scenario(repo_path)
assert inc_graph == clean_graph, "Incremental graph diverges from clean index"

```

This test suite prevented regression in issue #532 and runs across multiple languages (C++, Rust, Java, Python).

## Performance Characteristics for Large Codebases

The incremental design delivers measurable advantages that compound as repositories scale.

| Metric | Full Re-Index | Incremental Update | Improvement |
|--------|-------------|-------------------|-------------|
| Typical runtime | Minutes | Seconds | 10-100x |
| Memory allocation | All file ASTs | Changed files only | Bounded by edit size |
| CPU utilization | 100% parse workload | Proportional to changes | Minimal for small edits |
| Graph consistency | Guaranteed | Verified by test suite | Equivalent correctness |

The **deterministic graph** property is critical: after any sequence of incremental updates, the graph matches what a fresh full-index would produce. This eliminates "drift" that accumulates in systems with approximate incremental logic.

## Command-Line and Programmatic Interfaces

The system exposes incremental parsing through both CLI and Python APIs.

**CLI usage:**

```bash
cgr update_repository --path /path/to/repo --incremental

```

**Python programmatic interface:**

```python
from codebase_rag.graph_updater import GraphUpdater
from codebase_rag.storage import GraphStorage

store = GraphStorage("/tmp/cgr-store")
graph = store.load()

updater = GraphUpdater(store, incremental=True)
updater.update()

# Graph now reflects latest source state

```

Both interfaces accept the same `incremental` flag, ensuring consistent behavior across usage patterns.

## Key Implementation Files

| File | Purpose |
|------|---------|
| [`codebase_rag/graph_updater.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/graph_updater.py) | Core incremental engine with `GraphUpdater` class and re-hydration methods |
| [`codebase_rag/ast_cache.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/ast_cache.py) | Persistent AST storage for unchanged files |
| [`evals/incremental.py`](https://github.com/vitali87/code-graph-rag/blob/main/evals/incremental.py) | Correctness validation harness |
| `codebase_rag/tests/test_incremental_*.py` | Edge case coverage for inheritance, cross-file calls, and polyglot scenarios |
| `codebase_rag/parsers/*` | Language-specific parsing implementations |

## Summary

- **code-graph-rag** implements incremental parsing through coordinated change detection, selective re-parsing, AST caching, and graph re-hydration.
- The `GraphUpdater` class orchestrates the complete workflow in [`codebase_rag/graph_updater.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/graph_updater.py).
- File-hash comparison in `_update_file` provides precise, content-based change detection.
- `ASTCache` eliminates parsing overhead for unchanged files.
- Re-hydration methods restore derived relations (`_rehydrate_class_inheritance_from_graph`, `_rehydrate_call_edges`) to maintain graph completeness.
- The [`evals/incremental.py`](https://github.com/vitali87/code-graph-rag/blob/main/evals/incremental.py) test suite guarantees that incremental results match clean full-index results.

## Frequently Asked Questions

### How does code-graph-rag detect which files changed?

The system stores a SHA-256 hash of each file's content as a node property in the graph. On every incremental run, it computes fresh hashes and compares them against stored values. This content-based approach catches changes that file timestamps miss, including modifications that preserve modification time and git operations that alter content.

### What happens to cross-file relationships when only one file changes?

Cross-file edges (class inheritance, function calls, imports) are **purged before incremental parsing** and **re-hydrated afterward**. The `_rehydrate_class_inheritance_from_graph` and related methods traverse the persisted graph state to reconstruct these relationships from the complete file set. This ensures that changes in one file correctly propagate to edges involving unchanged files.

### When should I use incremental mode versus full re-indexing?

Use **incremental mode** for routine development workflows where small changes accumulate over time. Perform a **full re-index** when the storage format changes, after manual graph editing, or when diagnostic tools suggest graph corruption. The [`evals/incremental.py`](https://github.com/vitali87/code-graph-rag/blob/main/evals/incremental.py) harness can verify that incremental updates remain consistent with fresh indexes.