How Code-Graph-RAG Builds a Knowledge Graph from Source Code
Code-Graph-RAG constructs a queryable knowledge graph by parsing source files into ASTs, extracting definitions and call relationships, and streaming them into a graph database through a coordinated five-stage pipeline orchestrated by the GraphUpdater class.
Code-Graph-RAG transforms raw codebases into structured knowledge graphs that power semantic search and retrieval-augmented generation workflows. The system analyzes definitions, imports, and call relationships across multi-language repositories using Tree-Sitter parsers and language-specific front-ends. This article breaks down exactly how the vitali87/code-graph-rag repository implements this pipeline, from initial file walking to incremental updates.
Stage 1: Repository Walking and AST Parsing
The construction process begins in codebase_rag/graph_updater.py, where the GraphUpdater class initializes the parsing infrastructure. During initialization (GraphUpdater.__init__), the system creates a parser map that routes files to language-specific Tree-Sitter parsers, with fallback tiers including AstGrepTier and DocumentTier for unsupported languages.
The walk_eligible_files method traverses the repository, selecting appropriate parsers for each file extension. Each parsed file generates an Abstract Syntax Tree (AST) that is cached in the BoundedASTCache to minimize memory pressure during large codebase ingestion. This caching layer ensures that subsequent incremental runs can reuse previously parsed trees without reprocessing unchanged files.
Stage 2: Definition and Relationship Extraction
Once ASTs are available, the ProcessorFactory drives the extraction of graph entities. For every AST node, the factory emits definition nodes representing functions, classes, constants, and other symbols, while simultaneously recording import and call edges that map dependencies between these entities.
Language-specific front-ends enrich the raw Tree-Sitter data with additional semantic facts. The system implements dedicated methods such as _run_cpp_frontend and _run_csharp_frontend in GraphUpdater to handle language peculiarities like C++ macro expansions, Roslyn compiler facts for C#, or Go type information. These front-ends ensure the knowledge graph captures language-specific nuances that generic parsing might miss.
Stage 3: Streaming to Graph Storage
Extracted nodes and relationships flow through the _sink method to an Ingestor responsible for database writes. By default, Code-Graph-RAG uses the FilteringIngestor pipeline, which feeds into MemgraphIngestor to write data into a Memgraph instance via Bolt protocol.
The ingestor performs atomic writes, ensuring that either all entities from a file commit successfully or none do, maintaining graph consistency. During this phase, the system also updates parser fingerprints and metadata that enable incremental change detection in future runs.
Stage 4: Persistence and Graph Loading
After successful ingestion, the knowledge graph persists to a JSON file (defaulting to graph_code_index.pb or a user-specified path). This export captures the complete graph state including node properties, relationship types, and indexing metadata.
The codebase_rag/graph_loader.py module provides the GraphLoader class to reconstruct the graph from these exports. The loader rebuilds in-memory indices and offers query helpers such as find_nodes_by_label and summary methods, enabling offline analysis and rapid graph inspection without requiring a live database connection.
Stage 5: Incremental Re-ingestion for Large Repositories
For repositories that evolve over time, Code-Graph-RAG implements efficient incremental updates rather than full rebuilds. The system maintains hash caches (managed via _publish_hash_cache and _trustworthy_cache_mtime) that track file contents and modification times.
On subsequent runs, the updater compares current file hashes against cached values, reparsing only changed files while preserving the existing graph structure for unchanged portions of the codebase. This approach dramatically reduces indexing time for large repositories, enabling near real-time graph synchronization with the underlying code.
Implementing the Pipeline: Code Examples
Indexing a Repository into Memgraph
from codebase_rag.services.graph_service import MemgraphIngestor
from codebase_rag.graph_updater import GraphUpdater
from codebase_rag.parser_loader import load_parsers, load_queries
from pathlib import Path
# Load language-specific parsers and accompanying queries
parsers = load_parsers()
queries = load_queries()
# Create an ingestor that writes to a local Memgraph instance
ingestor = MemgraphIngestor(
uri="bolt://localhost:7687",
user="memgraph",
password="******"
)
# Initialize the updater for the target checkout
updater = GraphUpdater(
ingestor=ingestor,
repo_path=Path("/path/to/your/project"),
parsers=parsers,
queries=queries,
)
# Run a full build (re-ingest everything)
updater.reingest()
This example demonstrates how GraphUpdater coordinates the complete pipeline, from walking the repository to streaming extracted entities into Memgraph.
Loading and Querying a Persisted Graph
from codebase_rag.graph_loader import load_graph
# Load a previously exported JSON graph
loader = load_graph("graph_export.json")
# List how many nodes of each label exist
print(loader.summary().node_labels)
# Find the first 5 function nodes
functions = loader.find_nodes_by_label("Function")[:5]
for fn in functions:
print(fn.properties["qualified_name"])
The GraphLoader rebuilds fast lookup tables from the JSON export, providing convenient methods to inspect the knowledge graph without maintaining a live database connection.
Key Source Files and Architecture
-
codebase_rag/graph_updater.py– The core orchestrator containingGraphUpdater, which manages the entire pipeline from parsing through ingestion. Implements__init__,walk_eligible_files, language-specific front-ends (_run_cpp_frontend, etc.), and_sinkfor database writes. -
codebase_rag/graph_loader.py– Handles persistence and reloading via theGraphLoaderclass. Providesload,summary, andfind_nodes_by_labelmethods for querying exported graphs. -
codebase_rag/services/graph_service.py– ImplementsMemgraphIngestorand other backend ingestors that perform the actual atomic writes to graph databases. -
codebase_rag/parser_loader.py– Loads Tree-Sitter parsers and language-specific query files used by the updater to extract graph entities from source code. -
codebase_rag/parsers/– Directory containing language-specific parsing tiers including Tree-Sitter implementations, ast-grep fallbacks, and front-end modules likecpp_frontend.pythat add language-specific semantic facts. -
codebase_rag/services/graph_diff.py– Helper utilities for incremental updates, managing hash caches and comparing cached states with the current filesystem to minimize reprocessing.
Summary
- Code-Graph-RAG builds knowledge graphs through a five-stage pipeline: repository walking, AST parsing, entity extraction, graph streaming, and persistence.
- The
GraphUpdaterclass ingraph_updater.pyserves as the central orchestrator, managing parsers, language front-ends, and database ingestors. - Tree-Sitter provides the foundation for parsing, while language-specific front-ends (C++, C#, Go, Java) enrich the graph with semantic details like macros and type information.
- Incremental updates via hash caching enable efficient re-indexing of large repositories by processing only changed files.
- The system supports both live database ingestion (Memgraph) and offline JSON exports (
graph_code_index.pb) for flexible deployment scenarios.
Frequently Asked Questions
What parsing engine does Code-Graph-RAG use to analyze source code?
Code-Graph-RAG uses Tree-Sitter as its primary parsing engine, with parser_loader.py dynamically loading language-specific grammars and queries. For unsupported languages or fallback scenarios, the system employs ast-grep (AstGrepTier) or generic document parsing (DocumentTier) to ensure broad language coverage while maintaining AST precision where possible.
How does the system handle updates when code changes?
The system implements incremental re-ingestion using hash caching mechanisms defined in graph_diff.py and _publish_hash_cache. By comparing file hashes and modification times against cached values, GraphUpdater reprocesses only modified files while preserving the existing graph structure for unchanged code, significantly reducing indexing time for large repositories.
Can I export the knowledge graph without using a database?
Yes. While MemgraphIngestor provides live database streaming, the system can persist the complete knowledge graph to a JSON file (default graph_code_index.pb) using the export functionality in graph_loader.py. The GraphLoader class can later reload this file, rebuild in-memory indices, and provide full query capabilities without requiring a database connection.
Which programming languages receive specialized front-end processing?
Code-Graph-RAG provides specialized language front-ends for C++, C#, Go, Java, and other major languages through dedicated methods like _run_cpp_frontend and _run_csharp_frontend. These front-ends enrich raw Tree-Sitter ASTs with additional semantic facts such as C++ macro expansions, Roslyn compiler metadata for C#, and Go-specific type information.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →