# How code-graph-rag Parses Code for Graph Generation: Tree-Sitter AST to Knowledge Graph Pipeline

> Explore how code-graph-rag parses code for graph generation. It transforms Tree-Sitter ASTs into knowledge graphs using language analyzers and the GraphUpdater class. Learn the pipeline.

- Repository: [Vitali Avagyan/code-graph-rag](https://github.com/vitali87/code-graph-rag)
- Tags: internals
- Published: 2026-08-18

---

**code-graph-rag builds code graphs by loading Tree-Sitter language grammars, traversing abstract syntax trees with language-specific analyzers, and converting extracted entities into typed nodes and relationships via the GraphUpdater class.**

The open-source repository `vitali87/code-graph-rag` transforms raw source files into navigable knowledge graphs designed for retrieval-augmented generation (RAG) over codebases. Understanding how code-graph-rag parses code for graph generation requires examining its modular pipeline—from grammar loading and AST traversal to graph construction and persistence.

## Loading Tree-Sitter Language Grammars

Before parsing begins, `code-graph-rag` initializes language support through [`codebase_rag/tools/language.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/tools/language.py). This module discovers compiled Tree-Sitter language libraries (such as *tree_sitter_python*, *tree_sitter_javascript*, and *tree_sitter_rust*) and instantiates `Language` objects for each supported grammar.

These language objects are cached to eliminate redundant binary loading, ensuring that subsequent parsing operations reuse the same grammar instances without filesystem overhead.

```python
from codebase_rag.tools.language import get_language

# Load Python grammar once and cache for reuse

python_lang = get_language("python")  # Returns tree_sitter.Language instance

```

## Language-Specific AST Parsing

The system delegates actual AST analysis to specialized modules under `codebase_rag/parsers/`, with each language implementing its own extraction logic.

### Python AST Analysis

For Python files, [`codebase_rag/parsers/py/ast_analyzer.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/parsers/py/ast_analyzer.py) implements the `PythonAstAnalyzerMixin` class. This analyzer walks the Python AST to collect assignments, comprehensions, for-loops, return statements, and identifier resolutions. It executes cached Tree-Sitter queries (including `_PY_TRAVERSE_QUERY` and `cs.PY_RETURN_QUERY`) to locate nodes efficiently without redundant traversal.

The mixin provides `_traverse_single_pass`, a unified method that iterates over the AST using either a `QueryCursor` for targeted searches or a manual depth-first walk when queries fail. This ensures comprehensive coverage of class definitions, function signatures, and variable assignments while feeding data into the type-inference engine.

### JavaScript and TypeScript Extraction

JavaScript and TypeScript parsing utilizes [`codebase_rag/parsers/js_ts/utils.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/parsers/js_ts/utils.py), which supplies helpers like `find_js_method_in_ast` to locate method definitions within JS/TS AST structures. These utilities handle the specific syntactic patterns of ECMAScript-based languages, extracting function declarations, class methods, and import relationships.

### Rust Type Inference

For Rust code, [`codebase_rag/parsers/rs/type_inference.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/parsers/rs/type_inference.py) and companion modules import Tree-Sitter `Node` objects to extract function signatures, trait implementations, and type annotations. The Rust parser handles the language's ownership semantics and generic constraints while mapping syntactic constructs to graph-compatible entities.

## Unified AST Traversal Logic

Regardless of language, the parsing workflow follows a consistent pattern through analyzer mixins. The `_traverse_single_pass` method serves as the central traversal mechanism, accepting a root AST node and producing structured data about code entities.

This method implements fallback logic: it attempts optimized Tree-Sitter queries first, then resorts to manual recursive descent when query execution fails or returns incomplete results. Every relevant node—whether representing assignments, method definitions, or class declarations—gets visited and catalogued for relationship mapping.

## Converting AST Data to Graph Structures

Once the AST analyzers extract entities, [`codebase_rag/graph_updater.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/graph_updater.py) transforms this raw syntax data into formal graph components. The `GraphUpdater` class implements `update_graph_from_ast()`, which processes the parsed tree and creates:

- **GraphNodeRecord** objects representing functions, classes, modules, and variables
- **GraphRelRecord** objects defining typed relationships between nodes

The updater establishes specific edge types including:

- **CALLS** – Links a function to the functions it invokes
- **INHERITS** – Connects a class to its parent classes
- **IMPORTS** – Maps modules to their external dependencies
- **DEFINE** – Associates modules with entities they declare

During construction, `GraphUpdater` deduplicates nodes, resolves fully-qualified names, and maintains an in-memory graph structure ready for serialization.

## Persisting and Loading Code Graphs

After graph construction, [`codebase_rag/graph_loader.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/graph_loader.py) handles persistence through the `GraphLoader` class. The loader serializes the graph to JSON format containing the complete collection of `GraphNodeRecord` and `GraphRelRecord` objects, or deserializes existing graphs back into memory.

The loader provides query helpers including `find_nodes_by_label` and `get_node_by_id` to facilitate downstream RAG retrieval operations.

```python
from codebase_rag.graph_loader import GraphLoader, load_graph

# Serialize built graph to disk

loader = GraphLoader("/tmp/code_graph.json")
loader.save_graph(graph)

# Reload for later analysis

loaded_graph = load_graph("/tmp/code_graph.json")

```

## Complete Parsing-to-Graph Pipeline

The following example demonstrates the end-to-end workflow from source file to serialized graph:

```python
from pathlib import Path
from tree_sitter import Parser
from codebase_rag.tools.language import get_language
from codebase_rag.parsers.py.ast_analyzer import PythonAstAnalyzerMixin
from codebase_rag.graph_updater import GraphUpdater
from codebase_rag.graph_loader import GraphLoader

# 1. Initialize Tree-Sitter parser with language grammar

python_lang = get_language("python")
parser = Parser()
parser.set_language(python_lang)

# 2. Parse source file into AST

source_code = Path("my_module.py").read_bytes()
tree = parser.parse(source_code)

# 3. Extract entities via AST analyzer

class MyAnalyzer(PythonAstAnalyzerMixin):
    pass

analyzer = MyAnalyzer()
root_node = tree.root_node
comps, loops = analyzer._traverse_single_pass(root_node, {}, "my_module")

# 4. Build knowledge graph from AST

updater = GraphUpdater()
updater.update_graph_from_ast(
    root_node, 
    language="python", 
    module_qn="my_module"
)
graph = updater.graph  # Contains nodes and relationships

# 5. Persist graph for RAG queries

loader = GraphLoader("/tmp/my_graph.json")
loader.save_graph(graph)

```

## Summary

- **Grammar Initialization**: [`codebase_rag/tools/language.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/tools/language.py) loads and caches Tree-Sitter language binaries to avoid repeated filesystem operations.
- **Language Parsers**: Specialized modules under `codebase_rag/parsers/` (Python, JS/TS, Rust) extract language-specific AST information using targeted Tree-Sitter queries.
- **Traversal Strategy**: The `PythonAstAnalyzerMixin` and analogous classes implement `_traverse_single_pass` to ensure complete AST coverage via query-first or manual walk strategies.
- **Graph Construction**: `GraphUpdater.update_graph_from_ast()` converts AST entities into typed nodes and relationships (CALLS, INHERITS, IMPORTS, DEFINE).
- **Persistence**: `GraphLoader` serializes graphs to JSON and provides lookup utilities for downstream RAG consumption.

## Frequently Asked Questions

### What parsing technology does code-graph-rag use for multi-language support?

**code-graph-rag uses Tree-Sitter** as its foundation for multi-language parsing. According to the source code in [`codebase_rag/tools/language.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/tools/language.py), the system discovers and loads compiled Tree-Sitter language grammars (like *tree_sitter_python* and *tree_sitter_javascript*) to create reusable `Language` objects that power the AST parsers.

### How does the system handle different AST structures across programming languages?

The architecture delegates parsing to **language-specific modules** under `codebase_rag/parsers/`. Each language implements its own analyzer (such as `PythonAstAnalyzerMixin` for Python or [`type_inference.py`](https://github.com/vitali87/code-graph-rag/blob/main/type_inference.py) for Rust) that understands the specific AST node types of that language. These analyzers use cached Tree-Sitter queries to extract relevant syntactic entities while maintaining a consistent interface for the graph construction phase.

### What types of relationships are captured in the generated code graph?

The `GraphUpdater` class creates several relationship types including **CALLS** (function invocations), **INHERITS** (class inheritance), **IMPORTS** (module dependencies), and **DEFINE** (module-to-entity declarations). These relationships are stored as `GraphRelRecord` objects alongside `GraphNodeRecord` entities to create a fully navigable graph structure suitable for dependency analysis and RAG queries.

### Can the generated graph be saved and reused across sessions?

Yes, the `GraphLoader` class in [`codebase_rag/graph_loader.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/graph_loader.py) provides serialization capabilities. After construction via `GraphUpdater`, graphs can be persisted to JSON format using `loader.save_graph()`, then reloaded into memory later using `load_graph()`. The loader includes query methods like `find_nodes_by_label` to support efficient graph traversal during RAG operations.