How to Parse a Codebase and Update the Graph in Code-Graph-RAG

Code-Graph-RAG uses a three-pass pipeline orchestrated by GraphUpdater to parse source files into ASTs, extract definitions and references, and stream the resulting nodes and edges to a graph database, with support for incremental re-ingest when files change.

Code-Graph-RAG builds a Neo4j-style property graph representing the logical structure of software repositories. To parse a codebase and update the graph in Code-Graph-RAG, you use the GraphUpdater class located in codebase_rag/graph_updater.py, which orchestrates Tree-Sitter parsing, semantic analysis, and batched ingestion while maintaining auxiliary caches for performance.

The GraphUpdater Orchestrator

The GraphUpdater class serves as the central orchestrator for transforming source code into a queryable graph structure. It manages the entire lifecycle from file discovery through relationship resolution.

Initialization and Configuration

When instantiating GraphUpdater, you provide an ingestor (the database writer), a repo path, and language-specific parsers. The __init__ method in codebase_rag/graph_updater.py initializes several critical components:

  • FunctionRegistryTrie (function_registry.py) – Enables quick qualified-name lookup for cross-file call resolution
  • BoundedASTCache (ast_cache.py) – Lazily loads and evicts syntax trees to minimize memory pressure
  • ProcessorFactory (parsers/factory.py) – Builds the three processing tiers (definition, reference, and finding)
from pathlib import Path
from codebase_rag.graph_updater import GraphUpdater
from codebase_rag.parser_loader import load_parsers

# Load Tree-Sitter parsers and compiled queries

parsers, queries = load_parsers()

# Initialize the updater

updater = GraphUpdater(
    ingestor=ingestor,                    # Neo4j or mock ingestor

    repo_path=Path("/path/to/repo"),
    parsers=parsers,
    queries=queries,
    exclude={".git", "node_modules"},     # Optional path exclusions

    project_name="my_project"
)

Optional Language-Specific Front-Ends

Before Tree-Sitter processing begins, GraphUpdater can invoke specialized front-ends for deep semantic analysis. These are controlled via settings flags and run through private methods like _run_cpp_frontend, _run_csharp_frontend, and _run_go_frontend:

  • C++ – libclang frontend handles macro expansion and include edges
  • C# – Roslyn frontend captures inheritance hierarchies and LINQ call patterns
  • Go – go/types frontend provides accurate call target resolution

Results from these front-ends populate attributes such as _pending_cpp_macro_calls for merging during Pass-3.

The Three-Pass Parsing Pipeline

The core ingestion logic runs three sequential passes over the repository, implemented in GraphUpdater._process_files and related methods.

Pass 1 – File Discovery and AST Generation

The pipeline begins by walking the repository using walk_eligible_files from codebase_rag/utils/path_utils.py, respecting exclude/unignore filters. For each eligible file:

  1. Language detection occurs via get_language_for_extension in language_spec.py
  2. Tree-Sitter parsing generates the syntax tree using the appropriate Parser.parse call
  3. AST caching stores the tree in self.ast_cache for reuse in subsequent passes

The method records each (Path, SupportedLanguage) tuple in _parsed_files to preserve parsing order for incremental runs.

Pass 2 – Definition Extraction

Each cached AST is processed by the definition processor (self.factory.definition_processor). This phase extracts:

  • Structural nodes – Modules, packages, classes, functions, and methods as GraphNode objects (models.py)
  • Import relationships – ImportEdge objects representing module dependencies
  • Registry population – Qualified names are registered in FunctionRegistryTrie for cross-file resolution

If hybrid front-ends are active (e.g., C++ libclang), macro-generated nodes are injected here and pending macro-call information is staged for the next phase.

Pass 3 – Reference and Call Resolution

The reference processor (self.factory.reference_processor) traverses the ASTs again to resolve symbolic relationships:

  • Call site resolution matches caller/callee pairs using the FunctionRegistryTrie built in Pass-2
  • Call edges are generated as GraphRelationship objects with type CALL
  • Semantic fact merging applies external front-end data via _apply_semantic_facts and _apply_go_semantic_facts

At completion, the full set of GraphNode and GraphRelationship objects streams to the ingestor for persistence.

Incremental Re-Ingest for Changed Files

When only specific files change, GraphUpdater.reingest(paths, deleted=...) performs scoped updates without rebuilding the entire graph:

  1. Deletion – Removes stale sub-graphs for affected files via _reingest_delete
  2. Re-parsing – Processes only changed or new files through the three-pass pipeline
  3. Cache updates – Refreshes hash caches, directory modification times, and parser fingerprints to optimize subsequent runs

# After modifying specific files

changed_files = ["src/module/foo.py", "src/utils/helpers.py"]
deleted_files = ["src/module/old.py"]

updater.reingest(changed_files, deleted=deleted_files)

Loading and Querying the Graph with GraphLoader

After ingestion, GraphLoader (codebase_rag/graph_loader.py) reads the persisted JSON graph (typically graph.json) and provides fast lookup helpers:

from codebase_rag.graph_loader import GraphLoader

loader = GraphLoader("graph.json")
loader.load()

# Query by label

functions = loader.find_nodes_by_label("FUNCTION")

# Direct ID lookup

node = loader._nodes_by_id[42]

The loader constructs three indexes during load():

  • _nodes_by_id – Numeric ID to node mapping
  • _nodes_by_label – Grouping by NodeLabel enum values
  • Property indexes – Lazily built on first find_node_by_property call

Complete Implementation Example

Here is the full workflow from initialization through incremental updates:

from pathlib import Path
from codebase_rag.graph_updater import GraphUpdater
from codebase_rag.graph_loader import GraphLoader
from codebase_rag.ingestor import Ingestor
from codebase_rag.parser_loader import load_parsers

# 1. Configure the storage backend

ingestor = Ingestor(uri="bolt://localhost:7687", auth=("neo4j", "password"))

# 2. Load Tree-Sitter parsers and query objects

parsers, queries = load_parsers()

# 3. Initialize updater and run full index

updater = GraphUpdater(
    ingestor=ingestor,
    repo_path=Path("/path/to/repo"),
    parsers=parsers,
    queries=queries,
)
updater.run()  # Executes passes 1-3

# 4. Incremental update after changes

updater.reingest(
    paths=["src/core/engine.py"],
    deleted=["src/legacy/module.py"]
)

# 5. Load persisted graph for analysis

loader = GraphLoader("graph.json")
loader.load()
print(f"Total functions: {len(loader.find_nodes_by_label('FUNCTION'))}")

Summary

  • GraphUpdater (codebase_rag/graph_updater.py) orchestrates the three-pass parsing pipeline: file discovery, definition extraction, and reference resolution.
  • Tree-Sitter parsers handle the initial AST generation, while optional front-ends (C++, C#, Go) provide deep semantic analysis for specific languages.
  • Incremental re-ingest via reingest() updates only changed files, preserving cache consistency through hash and timestamp tracking.
  • GraphLoader (codebase_rag/graph_loader.py) provides fast read-only access to persisted graphs via multiple indexing strategies.
  • The system maintains auxiliary structures including FunctionRegistryTrie for qualified name resolution and BoundedASTCache for memory-efficient tree reuse.

Frequently Asked Questions

What file formats and languages does Code-Graph-RAG support?

Code-Graph-RAG uses Tree-Sitter for parsing, which supports all languages with Tree-Sitter grammars including Python, JavaScript, TypeScript, Java, C, C++, C#, Go, and Rust. Language detection occurs in language_spec.py via get_language_for_extension. Optional hybrid front-ends for C++ (libclang), C# (Roslyn), and Go (go/types) provide additional semantic details when toolchains are available.

How does incremental re-ingest maintain graph consistency?

The reingest method in codebase_rag/graph_updater.py first deletes existing nodes and edges for affected files using _reingest_delete, then re-runs the three-pass pipeline only for the specified paths. It updates auxiliary caches including file hashes, directory mtimes, and parser fingerprints to ensure subsequent runs correctly identify whether files require reprocessing.

What is the difference between GraphUpdater and GraphLoader?

GraphUpdater is a write-heavy orchestrator that parses source code, resolves symbols, and streams data to a database backend. GraphLoader is a read-only utility that loads a previously exported JSON graph file (graph.json) and builds in-memory indexes (_nodes_by_id, _nodes_by_label) for fast query-time access without reparsing source files.

How are language-specific features like C++ macros handled?

When settings.CPP_FRONTEND is enabled, GraphUpdater invokes _run_cpp_frontend before the main parsing passes to capture macro expansions and include relationships using libclang. These semantic facts populate _pending_cpp_macro_calls and similar attributes, which are merged into the graph during Pass-3 via _apply_semantic_facts, ensuring macro-generated code appears correctly in the final graph structure.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →