# How to Parse a Codebase and Update the Graph in Code-Graph-RAG

> Learn how to parse a codebase and update the graph in Code-Graph-RAG. Discover the three-pass pipeline that efficiently ingests source files into a graph database.

- Repository: [Vitali Avagyan/code-graph-rag](https://github.com/vitali87/code-graph-rag)
- Tags: how-to-guide
- Published: 2026-09-08

---

**Code-Graph-RAG uses a three-pass pipeline orchestrated by `GraphUpdater` to parse source files into ASTs, extract definitions and references, and stream the resulting nodes and edges to a graph database, with support for incremental re-ingest when files change.**

Code-Graph-RAG builds a Neo4j-style property graph representing the logical structure of software repositories. To parse a codebase and update the graph in Code-Graph-RAG, you use the `GraphUpdater` class located in [`codebase_rag/graph_updater.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/graph_updater.py), which orchestrates Tree-Sitter parsing, semantic analysis, and batched ingestion while maintaining auxiliary caches for performance.

## The GraphUpdater Orchestrator

The `GraphUpdater` class serves as the central orchestrator for transforming source code into a queryable graph structure. It manages the entire lifecycle from file discovery through relationship resolution.

### Initialization and Configuration

When instantiating `GraphUpdater`, you provide an **ingestor** (the database writer), a **repo path**, and language-specific parsers. The `__init__` method in [`codebase_rag/graph_updater.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/graph_updater.py) initializes several critical components:

- **`FunctionRegistryTrie`** ([`function_registry.py`](https://github.com/vitali87/code-graph-rag/blob/main/function_registry.py)) – Enables quick qualified-name lookup for cross-file call resolution
- **`BoundedASTCache`** ([`ast_cache.py`](https://github.com/vitali87/code-graph-rag/blob/main/ast_cache.py)) – Lazily loads and evicts syntax trees to minimize memory pressure
- **`ProcessorFactory`** ([`parsers/factory.py`](https://github.com/vitali87/code-graph-rag/blob/main/parsers/factory.py)) – Builds the three processing tiers (definition, reference, and finding)

```python
from pathlib import Path
from codebase_rag.graph_updater import GraphUpdater
from codebase_rag.parser_loader import load_parsers

# Load Tree-Sitter parsers and compiled queries

parsers, queries = load_parsers()

# Initialize the updater

updater = GraphUpdater(
    ingestor=ingestor,                    # Neo4j or mock ingestor

    repo_path=Path("/path/to/repo"),
    parsers=parsers,
    queries=queries,
    exclude={".git", "node_modules"},     # Optional path exclusions

    project_name="my_project"
)

```

### Optional Language-Specific Front-Ends

Before Tree-Sitter processing begins, `GraphUpdater` can invoke specialized front-ends for deep semantic analysis. These are controlled via settings flags and run through private methods like `_run_cpp_frontend`, `_run_csharp_frontend`, and `_run_go_frontend`:

- **C++** – `libclang` frontend handles macro expansion and include edges
- **C#** – Roslyn frontend captures inheritance hierarchies and LINQ call patterns  
- **Go** – `go/types` frontend provides accurate call target resolution

Results from these front-ends populate attributes such as `_pending_cpp_macro_calls` for merging during Pass-3.

## The Three-Pass Parsing Pipeline

The core ingestion logic runs three sequential passes over the repository, implemented in `GraphUpdater._process_files` and related methods.

### Pass 1 – File Discovery and AST Generation

The pipeline begins by walking the repository using `walk_eligible_files` from [`codebase_rag/utils/path_utils.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/utils/path_utils.py), respecting exclude/unignore filters. For each eligible file:

1. **Language detection** occurs via `get_language_for_extension` in [`language_spec.py`](https://github.com/vitali87/code-graph-rag/blob/main/language_spec.py)
2. **Tree-Sitter parsing** generates the syntax tree using the appropriate `Parser.parse` call
3. **AST caching** stores the tree in `self.ast_cache` for reuse in subsequent passes

The method records each `(Path, SupportedLanguage)` tuple in `_parsed_files` to preserve parsing order for incremental runs.

### Pass 2 – Definition Extraction

Each cached AST is processed by the **definition processor** (`self.factory.definition_processor`). This phase extracts:

- **Structural nodes** – Modules, packages, classes, functions, and methods as `GraphNode` objects ([`models.py`](https://github.com/vitali87/code-graph-rag/blob/main/models.py))
- **Import relationships** – `ImportEdge` objects representing module dependencies
- **Registry population** – Qualified names are registered in `FunctionRegistryTrie` for cross-file resolution

If hybrid front-ends are active (e.g., C++ `libclang`), macro-generated nodes are injected here and pending macro-call information is staged for the next phase.

### Pass 3 – Reference and Call Resolution

The **reference processor** (`self.factory.reference_processor`) traverses the ASTs again to resolve symbolic relationships:

- **Call site resolution** matches caller/callee pairs using the `FunctionRegistryTrie` built in Pass-2
- **Call edges** are generated as `GraphRelationship` objects with type `CALL`
- **Semantic fact merging** applies external front-end data via `_apply_semantic_facts` and `_apply_go_semantic_facts`

At completion, the full set of `GraphNode` and `GraphRelationship` objects streams to the ingestor for persistence.

## Incremental Re-Ingest for Changed Files

When only specific files change, `GraphUpdater.reingest(paths, deleted=...)` performs scoped updates without rebuilding the entire graph:

1. **Deletion** – Removes stale sub-graphs for affected files via `_reingest_delete`
2. **Re-parsing** – Processes only changed or new files through the three-pass pipeline
3. **Cache updates** – Refreshes hash caches, directory modification times, and parser fingerprints to optimize subsequent runs

```python

# After modifying specific files

changed_files = ["src/module/foo.py", "src/utils/helpers.py"]
deleted_files = ["src/module/old.py"]

updater.reingest(changed_files, deleted=deleted_files)

```

## Loading and Querying the Graph with GraphLoader

After ingestion, `GraphLoader` ([`codebase_rag/graph_loader.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/graph_loader.py)) reads the persisted JSON graph (typically [`graph.json`](https://github.com/vitali87/code-graph-rag/blob/main/graph.json)) and provides fast lookup helpers:

```python
from codebase_rag.graph_loader import GraphLoader

loader = GraphLoader("graph.json")
loader.load()

# Query by label

functions = loader.find_nodes_by_label("FUNCTION")

# Direct ID lookup

node = loader._nodes_by_id[42]

```

The loader constructs three indexes during `load()`:
- **`_nodes_by_id`** – Numeric ID to node mapping
- **`_nodes_by_label`** – Grouping by `NodeLabel` enum values  
- **Property indexes** – Lazily built on first `find_node_by_property` call

## Complete Implementation Example

Here is the full workflow from initialization through incremental updates:

```python
from pathlib import Path
from codebase_rag.graph_updater import GraphUpdater
from codebase_rag.graph_loader import GraphLoader
from codebase_rag.ingestor import Ingestor
from codebase_rag.parser_loader import load_parsers

# 1. Configure the storage backend

ingestor = Ingestor(uri="bolt://localhost:7687", auth=("neo4j", "password"))

# 2. Load Tree-Sitter parsers and query objects

parsers, queries = load_parsers()

# 3. Initialize updater and run full index

updater = GraphUpdater(
    ingestor=ingestor,
    repo_path=Path("/path/to/repo"),
    parsers=parsers,
    queries=queries,
)
updater.run()  # Executes passes 1-3

# 4. Incremental update after changes

updater.reingest(
    paths=["src/core/engine.py"],
    deleted=["src/legacy/module.py"]
)

# 5. Load persisted graph for analysis

loader = GraphLoader("graph.json")
loader.load()
print(f"Total functions: {len(loader.find_nodes_by_label('FUNCTION'))}")

```

## Summary

- **`GraphUpdater`** ([`codebase_rag/graph_updater.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/graph_updater.py)) orchestrates the three-pass parsing pipeline: file discovery, definition extraction, and reference resolution.
- **Tree-Sitter parsers** handle the initial AST generation, while optional front-ends (C++, C#, Go) provide deep semantic analysis for specific languages.
- **Incremental re-ingest** via `reingest()` updates only changed files, preserving cache consistency through hash and timestamp tracking.
- **`GraphLoader`** ([`codebase_rag/graph_loader.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/graph_loader.py)) provides fast read-only access to persisted graphs via multiple indexing strategies.
- The system maintains auxiliary structures including `FunctionRegistryTrie` for qualified name resolution and `BoundedASTCache` for memory-efficient tree reuse.

## Frequently Asked Questions

### What file formats and languages does Code-Graph-RAG support?

Code-Graph-RAG uses Tree-Sitter for parsing, which supports all languages with Tree-Sitter grammars including Python, JavaScript, TypeScript, Java, C, C++, C#, Go, and Rust. Language detection occurs in [`language_spec.py`](https://github.com/vitali87/code-graph-rag/blob/main/language_spec.py) via `get_language_for_extension`. Optional hybrid front-ends for C++ (libclang), C# (Roslyn), and Go (go/types) provide additional semantic details when toolchains are available.

### How does incremental re-ingest maintain graph consistency?

The `reingest` method in [`codebase_rag/graph_updater.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/graph_updater.py) first deletes existing nodes and edges for affected files using `_reingest_delete`, then re-runs the three-pass pipeline only for the specified paths. It updates auxiliary caches including file hashes, directory mtimes, and parser fingerprints to ensure subsequent runs correctly identify whether files require reprocessing.

### What is the difference between GraphUpdater and GraphLoader?

`GraphUpdater` is a write-heavy orchestrator that parses source code, resolves symbols, and streams data to a database backend. `GraphLoader` is a read-only utility that loads a previously exported JSON graph file ([`graph.json`](https://github.com/vitali87/code-graph-rag/blob/main/graph.json)) and builds in-memory indexes (`_nodes_by_id`, `_nodes_by_label`) for fast query-time access without reparsing source files.

### How are language-specific features like C++ macros handled?

When `settings.CPP_FRONTEND` is enabled, `GraphUpdater` invokes `_run_cpp_frontend` before the main parsing passes to capture macro expansions and include relationships using `libclang`. These semantic facts populate `_pending_cpp_macro_calls` and similar attributes, which are merged into the graph during Pass-3 via `_apply_semantic_facts`, ensuring macro-generated code appears correctly in the final graph structure.