How to Parse a Codebase and Update the Graph in Code-Graph-RAG
Code-Graph-RAG uses a three-pass pipeline orchestrated by GraphUpdater to parse source files into ASTs, extract definitions and references, and stream the resulting nodes and edges to a graph database, with support for incremental re-ingest when files change.
Code-Graph-RAG builds a Neo4j-style property graph representing the logical structure of software repositories. To parse a codebase and update the graph in Code-Graph-RAG, you use the GraphUpdater class located in codebase_rag/graph_updater.py, which orchestrates Tree-Sitter parsing, semantic analysis, and batched ingestion while maintaining auxiliary caches for performance.
The GraphUpdater Orchestrator
The GraphUpdater class serves as the central orchestrator for transforming source code into a queryable graph structure. It manages the entire lifecycle from file discovery through relationship resolution.
Initialization and Configuration
When instantiating GraphUpdater, you provide an ingestor (the database writer), a repo path, and language-specific parsers. The __init__ method in codebase_rag/graph_updater.py initializes several critical components:
FunctionRegistryTrie(function_registry.py) – Enables quick qualified-name lookup for cross-file call resolutionBoundedASTCache(ast_cache.py) – Lazily loads and evicts syntax trees to minimize memory pressureProcessorFactory(parsers/factory.py) – Builds the three processing tiers (definition, reference, and finding)
from pathlib import Path
from codebase_rag.graph_updater import GraphUpdater
from codebase_rag.parser_loader import load_parsers
# Load Tree-Sitter parsers and compiled queries
parsers, queries = load_parsers()
# Initialize the updater
updater = GraphUpdater(
ingestor=ingestor, # Neo4j or mock ingestor
repo_path=Path("/path/to/repo"),
parsers=parsers,
queries=queries,
exclude={".git", "node_modules"}, # Optional path exclusions
project_name="my_project"
)
Optional Language-Specific Front-Ends
Before Tree-Sitter processing begins, GraphUpdater can invoke specialized front-ends for deep semantic analysis. These are controlled via settings flags and run through private methods like _run_cpp_frontend, _run_csharp_frontend, and _run_go_frontend:
- C++ –
libclangfrontend handles macro expansion and include edges - C# – Roslyn frontend captures inheritance hierarchies and LINQ call patterns
- Go –
go/typesfrontend provides accurate call target resolution
Results from these front-ends populate attributes such as _pending_cpp_macro_calls for merging during Pass-3.
The Three-Pass Parsing Pipeline
The core ingestion logic runs three sequential passes over the repository, implemented in GraphUpdater._process_files and related methods.
Pass 1 – File Discovery and AST Generation
The pipeline begins by walking the repository using walk_eligible_files from codebase_rag/utils/path_utils.py, respecting exclude/unignore filters. For each eligible file:
- Language detection occurs via
get_language_for_extensioninlanguage_spec.py - Tree-Sitter parsing generates the syntax tree using the appropriate
Parser.parsecall - AST caching stores the tree in
self.ast_cachefor reuse in subsequent passes
The method records each (Path, SupportedLanguage) tuple in _parsed_files to preserve parsing order for incremental runs.
Pass 2 – Definition Extraction
Each cached AST is processed by the definition processor (self.factory.definition_processor). This phase extracts:
- Structural nodes – Modules, packages, classes, functions, and methods as
GraphNodeobjects (models.py) - Import relationships –
ImportEdgeobjects representing module dependencies - Registry population – Qualified names are registered in
FunctionRegistryTriefor cross-file resolution
If hybrid front-ends are active (e.g., C++ libclang), macro-generated nodes are injected here and pending macro-call information is staged for the next phase.
Pass 3 – Reference and Call Resolution
The reference processor (self.factory.reference_processor) traverses the ASTs again to resolve symbolic relationships:
- Call site resolution matches caller/callee pairs using the
FunctionRegistryTriebuilt in Pass-2 - Call edges are generated as
GraphRelationshipobjects with typeCALL - Semantic fact merging applies external front-end data via
_apply_semantic_factsand_apply_go_semantic_facts
At completion, the full set of GraphNode and GraphRelationship objects streams to the ingestor for persistence.
Incremental Re-Ingest for Changed Files
When only specific files change, GraphUpdater.reingest(paths, deleted=...) performs scoped updates without rebuilding the entire graph:
- Deletion – Removes stale sub-graphs for affected files via
_reingest_delete - Re-parsing – Processes only changed or new files through the three-pass pipeline
- Cache updates – Refreshes hash caches, directory modification times, and parser fingerprints to optimize subsequent runs
# After modifying specific files
changed_files = ["src/module/foo.py", "src/utils/helpers.py"]
deleted_files = ["src/module/old.py"]
updater.reingest(changed_files, deleted=deleted_files)
Loading and Querying the Graph with GraphLoader
After ingestion, GraphLoader (codebase_rag/graph_loader.py) reads the persisted JSON graph (typically graph.json) and provides fast lookup helpers:
from codebase_rag.graph_loader import GraphLoader
loader = GraphLoader("graph.json")
loader.load()
# Query by label
functions = loader.find_nodes_by_label("FUNCTION")
# Direct ID lookup
node = loader._nodes_by_id[42]
The loader constructs three indexes during load():
_nodes_by_id– Numeric ID to node mapping_nodes_by_label– Grouping byNodeLabelenum values- Property indexes – Lazily built on first
find_node_by_propertycall
Complete Implementation Example
Here is the full workflow from initialization through incremental updates:
from pathlib import Path
from codebase_rag.graph_updater import GraphUpdater
from codebase_rag.graph_loader import GraphLoader
from codebase_rag.ingestor import Ingestor
from codebase_rag.parser_loader import load_parsers
# 1. Configure the storage backend
ingestor = Ingestor(uri="bolt://localhost:7687", auth=("neo4j", "password"))
# 2. Load Tree-Sitter parsers and query objects
parsers, queries = load_parsers()
# 3. Initialize updater and run full index
updater = GraphUpdater(
ingestor=ingestor,
repo_path=Path("/path/to/repo"),
parsers=parsers,
queries=queries,
)
updater.run() # Executes passes 1-3
# 4. Incremental update after changes
updater.reingest(
paths=["src/core/engine.py"],
deleted=["src/legacy/module.py"]
)
# 5. Load persisted graph for analysis
loader = GraphLoader("graph.json")
loader.load()
print(f"Total functions: {len(loader.find_nodes_by_label('FUNCTION'))}")
Summary
GraphUpdater(codebase_rag/graph_updater.py) orchestrates the three-pass parsing pipeline: file discovery, definition extraction, and reference resolution.- Tree-Sitter parsers handle the initial AST generation, while optional front-ends (C++, C#, Go) provide deep semantic analysis for specific languages.
- Incremental re-ingest via
reingest()updates only changed files, preserving cache consistency through hash and timestamp tracking. GraphLoader(codebase_rag/graph_loader.py) provides fast read-only access to persisted graphs via multiple indexing strategies.- The system maintains auxiliary structures including
FunctionRegistryTriefor qualified name resolution andBoundedASTCachefor memory-efficient tree reuse.
Frequently Asked Questions
What file formats and languages does Code-Graph-RAG support?
Code-Graph-RAG uses Tree-Sitter for parsing, which supports all languages with Tree-Sitter grammars including Python, JavaScript, TypeScript, Java, C, C++, C#, Go, and Rust. Language detection occurs in language_spec.py via get_language_for_extension. Optional hybrid front-ends for C++ (libclang), C# (Roslyn), and Go (go/types) provide additional semantic details when toolchains are available.
How does incremental re-ingest maintain graph consistency?
The reingest method in codebase_rag/graph_updater.py first deletes existing nodes and edges for affected files using _reingest_delete, then re-runs the three-pass pipeline only for the specified paths. It updates auxiliary caches including file hashes, directory mtimes, and parser fingerprints to ensure subsequent runs correctly identify whether files require reprocessing.
What is the difference between GraphUpdater and GraphLoader?
GraphUpdater is a write-heavy orchestrator that parses source code, resolves symbols, and streams data to a database backend. GraphLoader is a read-only utility that loads a previously exported JSON graph file (graph.json) and builds in-memory indexes (_nodes_by_id, _nodes_by_label) for fast query-time access without reparsing source files.
How are language-specific features like C++ macros handled?
When settings.CPP_FRONTEND is enabled, GraphUpdater invokes _run_cpp_frontend before the main parsing passes to capture macro expansions and include relationships using libclang. These semantic facts populate _pending_cpp_macro_calls and similar attributes, which are merged into the graph during Pass-3 via _apply_semantic_facts, ensuring macro-generated code appears correctly in the final graph structure.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →