# How Code-Graph-RAG Builds a Knowledge Graph from Source Code

> Discover how Code-Graph-RAG builds a knowledge graph from source code. Learn about its five-stage pipeline for parsing, extraction, and graph database streaming.

- Repository: [Vitali Avagyan/code-graph-rag](https://github.com/vitali87/code-graph-rag)
- Tags: how-to-guide
- Published: 2026-09-08

---

**Code-Graph-RAG constructs a queryable knowledge graph by parsing source files into ASTs, extracting definitions and call relationships, and streaming them into a graph database through a coordinated five-stage pipeline orchestrated by the `GraphUpdater` class.**

Code-Graph-RAG transforms raw codebases into structured knowledge graphs that power semantic search and retrieval-augmented generation workflows. The system analyzes definitions, imports, and call relationships across multi-language repositories using Tree-Sitter parsers and language-specific front-ends. This article breaks down exactly how the `vitali87/code-graph-rag` repository implements this pipeline, from initial file walking to incremental updates.

## Stage 1: Repository Walking and AST Parsing

The construction process begins in [`codebase_rag/graph_updater.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/graph_updater.py), where the `GraphUpdater` class initializes the parsing infrastructure. During initialization (`GraphUpdater.__init__`), the system creates a parser map that routes files to language-specific **Tree-Sitter** parsers, with fallback tiers including **AstGrepTier** and **DocumentTier** for unsupported languages.

The `walk_eligible_files` method traverses the repository, selecting appropriate parsers for each file extension. Each parsed file generates an **Abstract Syntax Tree (AST)** that is cached in the `BoundedASTCache` to minimize memory pressure during large codebase ingestion. This caching layer ensures that subsequent incremental runs can reuse previously parsed trees without reprocessing unchanged files.

## Stage 2: Definition and Relationship Extraction

Once ASTs are available, the `ProcessorFactory` drives the extraction of graph entities. For every AST node, the factory emits **definition nodes** representing functions, classes, constants, and other symbols, while simultaneously recording **import** and **call** edges that map dependencies between these entities.

Language-specific front-ends enrich the raw Tree-Sitter data with additional semantic facts. The system implements dedicated methods such as `_run_cpp_frontend` and `_run_csharp_frontend` in `GraphUpdater` to handle language peculiarities like C++ macro expansions, Roslyn compiler facts for C#, or Go type information. These front-ends ensure the knowledge graph captures language-specific nuances that generic parsing might miss.

## Stage 3: Streaming to Graph Storage

Extracted nodes and relationships flow through the `_sink` method to an **Ingestor** responsible for database writes. By default, Code-Graph-RAG uses the `FilteringIngestor` pipeline, which feeds into `MemgraphIngestor` to write data into a Memgraph instance via Bolt protocol.

The ingestor performs **atomic writes**, ensuring that either all entities from a file commit successfully or none do, maintaining graph consistency. During this phase, the system also updates parser fingerprints and metadata that enable incremental change detection in future runs.

## Stage 4: Persistence and Graph Loading

After successful ingestion, the knowledge graph persists to a JSON file (defaulting to `graph_code_index.pb` or a user-specified path). This export captures the complete graph state including node properties, relationship types, and indexing metadata.

The [`codebase_rag/graph_loader.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/graph_loader.py) module provides the `GraphLoader` class to reconstruct the graph from these exports. The loader rebuilds in-memory indices and offers query helpers such as `find_nodes_by_label` and `summary` methods, enabling offline analysis and rapid graph inspection without requiring a live database connection.

## Stage 5: Incremental Re-ingestion for Large Repositories

For repositories that evolve over time, Code-Graph-RAG implements efficient incremental updates rather than full rebuilds. The system maintains hash caches (managed via `_publish_hash_cache` and `_trustworthy_cache_mtime`) that track file contents and modification times.

On subsequent runs, the updater compares current file hashes against cached values, reparsing only changed files while preserving the existing graph structure for unchanged portions of the codebase. This approach dramatically reduces indexing time for large repositories, enabling near real-time graph synchronization with the underlying code.

## Implementing the Pipeline: Code Examples

### Indexing a Repository into Memgraph

```python
from codebase_rag.services.graph_service import MemgraphIngestor
from codebase_rag.graph_updater import GraphUpdater
from codebase_rag.parser_loader import load_parsers, load_queries
from pathlib import Path

# Load language-specific parsers and accompanying queries

parsers = load_parsers()
queries = load_queries()

# Create an ingestor that writes to a local Memgraph instance

ingestor = MemgraphIngestor(
    uri="bolt://localhost:7687", 
    user="memgraph", 
    password="******"
)

# Initialize the updater for the target checkout

updater = GraphUpdater(
    ingestor=ingestor,
    repo_path=Path("/path/to/your/project"),
    parsers=parsers,
    queries=queries,
)

# Run a full build (re-ingest everything)

updater.reingest()

```

This example demonstrates how `GraphUpdater` coordinates the complete pipeline, from walking the repository to streaming extracted entities into Memgraph.

### Loading and Querying a Persisted Graph

```python
from codebase_rag.graph_loader import load_graph

# Load a previously exported JSON graph

loader = load_graph("graph_export.json")

# List how many nodes of each label exist

print(loader.summary().node_labels)

# Find the first 5 function nodes

functions = loader.find_nodes_by_label("Function")[:5]
for fn in functions:
    print(fn.properties["qualified_name"])

```

The `GraphLoader` rebuilds fast lookup tables from the JSON export, providing convenient methods to inspect the knowledge graph without maintaining a live database connection.

## Key Source Files and Architecture

- **[`codebase_rag/graph_updater.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/graph_updater.py)** – The core orchestrator containing `GraphUpdater`, which manages the entire pipeline from parsing through ingestion. Implements `__init__`, `walk_eligible_files`, language-specific front-ends (`_run_cpp_frontend`, etc.), and `_sink` for database writes.

- **[`codebase_rag/graph_loader.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/graph_loader.py)** – Handles persistence and reloading via the `GraphLoader` class. Provides `load`, `summary`, and `find_nodes_by_label` methods for querying exported graphs.

- **[`codebase_rag/services/graph_service.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/services/graph_service.py)** – Implements `MemgraphIngestor` and other backend ingestors that perform the actual atomic writes to graph databases.

- **[`codebase_rag/parser_loader.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/parser_loader.py)** – Loads Tree-Sitter parsers and language-specific query files used by the updater to extract graph entities from source code.

- **`codebase_rag/parsers/`** – Directory containing language-specific parsing tiers including Tree-Sitter implementations, ast-grep fallbacks, and front-end modules like [`cpp_frontend.py`](https://github.com/vitali87/code-graph-rag/blob/main/cpp_frontend.py) that add language-specific semantic facts.

- **[`codebase_rag/services/graph_diff.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/services/graph_diff.py)** – Helper utilities for incremental updates, managing hash caches and comparing cached states with the current filesystem to minimize reprocessing.

## Summary

- **Code-Graph-RAG** builds knowledge graphs through a five-stage pipeline: repository walking, AST parsing, entity extraction, graph streaming, and persistence.
- The `GraphUpdater` class in [`graph_updater.py`](https://github.com/vitali87/code-graph-rag/blob/main/graph_updater.py) serves as the central orchestrator, managing parsers, language front-ends, and database ingestors.
- **Tree-Sitter** provides the foundation for parsing, while language-specific front-ends (C++, C#, Go, Java) enrich the graph with semantic details like macros and type information.
- **Incremental updates** via hash caching enable efficient re-indexing of large repositories by processing only changed files.
- The system supports both live database ingestion (Memgraph) and offline JSON exports (`graph_code_index.pb`) for flexible deployment scenarios.

## Frequently Asked Questions

### What parsing engine does Code-Graph-RAG use to analyze source code?

Code-Graph-RAG uses **Tree-Sitter** as its primary parsing engine, with [`parser_loader.py`](https://github.com/vitali87/code-graph-rag/blob/main/parser_loader.py) dynamically loading language-specific grammars and queries. For unsupported languages or fallback scenarios, the system employs **ast-grep** (AstGrepTier) or generic document parsing (DocumentTier) to ensure broad language coverage while maintaining AST precision where possible.

### How does the system handle updates when code changes?

The system implements **incremental re-ingestion** using hash caching mechanisms defined in [`graph_diff.py`](https://github.com/vitali87/code-graph-rag/blob/main/graph_diff.py) and `_publish_hash_cache`. By comparing file hashes and modification times against cached values, `GraphUpdater` reprocesses only modified files while preserving the existing graph structure for unchanged code, significantly reducing indexing time for large repositories.

### Can I export the knowledge graph without using a database?

Yes. While `MemgraphIngestor` provides live database streaming, the system can persist the complete knowledge graph to a JSON file (default `graph_code_index.pb`) using the export functionality in [`graph_loader.py`](https://github.com/vitali87/code-graph-rag/blob/main/graph_loader.py). The `GraphLoader` class can later reload this file, rebuild in-memory indices, and provide full query capabilities without requiring a database connection.

### Which programming languages receive specialized front-end processing?

Code-Graph-RAG provides specialized language front-ends for **C++**, **C#**, **Go**, **Java**, and other major languages through dedicated methods like `_run_cpp_frontend` and `_run_csharp_frontend`. These front-ends enrich raw Tree-Sitter ASTs with additional semantic facts such as C++ macro expansions, Roslyn compiler metadata for C#, and Go-specific type information.