How code-graph-rag Parses Code for Graph Generation: Tree-Sitter AST to Knowledge Graph Pipeline

code-graph-rag builds code graphs by loading Tree-Sitter language grammars, traversing abstract syntax trees with language-specific analyzers, and converting extracted entities into typed nodes and relationships via the GraphUpdater class.

The open-source repository vitali87/code-graph-rag transforms raw source files into navigable knowledge graphs designed for retrieval-augmented generation (RAG) over codebases. Understanding how code-graph-rag parses code for graph generation requires examining its modular pipeline—from grammar loading and AST traversal to graph construction and persistence.

Loading Tree-Sitter Language Grammars

Before parsing begins, code-graph-rag initializes language support through codebase_rag/tools/language.py. This module discovers compiled Tree-Sitter language libraries (such as tree_sitter_python, tree_sitter_javascript, and tree_sitter_rust) and instantiates Language objects for each supported grammar.

These language objects are cached to eliminate redundant binary loading, ensuring that subsequent parsing operations reuse the same grammar instances without filesystem overhead.

from codebase_rag.tools.language import get_language

# Load Python grammar once and cache for reuse

python_lang = get_language("python")  # Returns tree_sitter.Language instance

Language-Specific AST Parsing

The system delegates actual AST analysis to specialized modules under codebase_rag/parsers/, with each language implementing its own extraction logic.

Python AST Analysis

For Python files, codebase_rag/parsers/py/ast_analyzer.py implements the PythonAstAnalyzerMixin class. This analyzer walks the Python AST to collect assignments, comprehensions, for-loops, return statements, and identifier resolutions. It executes cached Tree-Sitter queries (including _PY_TRAVERSE_QUERY and cs.PY_RETURN_QUERY) to locate nodes efficiently without redundant traversal.

The mixin provides _traverse_single_pass, a unified method that iterates over the AST using either a QueryCursor for targeted searches or a manual depth-first walk when queries fail. This ensures comprehensive coverage of class definitions, function signatures, and variable assignments while feeding data into the type-inference engine.

JavaScript and TypeScript Extraction

JavaScript and TypeScript parsing utilizes codebase_rag/parsers/js_ts/utils.py, which supplies helpers like find_js_method_in_ast to locate method definitions within JS/TS AST structures. These utilities handle the specific syntactic patterns of ECMAScript-based languages, extracting function declarations, class methods, and import relationships.

Rust Type Inference

For Rust code, codebase_rag/parsers/rs/type_inference.py and companion modules import Tree-Sitter Node objects to extract function signatures, trait implementations, and type annotations. The Rust parser handles the language's ownership semantics and generic constraints while mapping syntactic constructs to graph-compatible entities.

Unified AST Traversal Logic

Regardless of language, the parsing workflow follows a consistent pattern through analyzer mixins. The _traverse_single_pass method serves as the central traversal mechanism, accepting a root AST node and producing structured data about code entities.

This method implements fallback logic: it attempts optimized Tree-Sitter queries first, then resorts to manual recursive descent when query execution fails or returns incomplete results. Every relevant node—whether representing assignments, method definitions, or class declarations—gets visited and catalogued for relationship mapping.

Converting AST Data to Graph Structures

Once the AST analyzers extract entities, codebase_rag/graph_updater.py transforms this raw syntax data into formal graph components. The GraphUpdater class implements update_graph_from_ast(), which processes the parsed tree and creates:

  • GraphNodeRecord objects representing functions, classes, modules, and variables
  • GraphRelRecord objects defining typed relationships between nodes

The updater establishes specific edge types including:

  • CALLS – Links a function to the functions it invokes
  • INHERITS – Connects a class to its parent classes
  • IMPORTS – Maps modules to their external dependencies
  • DEFINE – Associates modules with entities they declare

During construction, GraphUpdater deduplicates nodes, resolves fully-qualified names, and maintains an in-memory graph structure ready for serialization.

Persisting and Loading Code Graphs

After graph construction, codebase_rag/graph_loader.py handles persistence through the GraphLoader class. The loader serializes the graph to JSON format containing the complete collection of GraphNodeRecord and GraphRelRecord objects, or deserializes existing graphs back into memory.

The loader provides query helpers including find_nodes_by_label and get_node_by_id to facilitate downstream RAG retrieval operations.

from codebase_rag.graph_loader import GraphLoader, load_graph

# Serialize built graph to disk

loader = GraphLoader("/tmp/code_graph.json")
loader.save_graph(graph)

# Reload for later analysis

loaded_graph = load_graph("/tmp/code_graph.json")

Complete Parsing-to-Graph Pipeline

The following example demonstrates the end-to-end workflow from source file to serialized graph:

from pathlib import Path
from tree_sitter import Parser
from codebase_rag.tools.language import get_language
from codebase_rag.parsers.py.ast_analyzer import PythonAstAnalyzerMixin
from codebase_rag.graph_updater import GraphUpdater
from codebase_rag.graph_loader import GraphLoader

# 1. Initialize Tree-Sitter parser with language grammar

python_lang = get_language("python")
parser = Parser()
parser.set_language(python_lang)

# 2. Parse source file into AST

source_code = Path("my_module.py").read_bytes()
tree = parser.parse(source_code)

# 3. Extract entities via AST analyzer

class MyAnalyzer(PythonAstAnalyzerMixin):
    pass

analyzer = MyAnalyzer()
root_node = tree.root_node
comps, loops = analyzer._traverse_single_pass(root_node, {}, "my_module")

# 4. Build knowledge graph from AST

updater = GraphUpdater()
updater.update_graph_from_ast(
    root_node, 
    language="python", 
    module_qn="my_module"
)
graph = updater.graph  # Contains nodes and relationships

# 5. Persist graph for RAG queries

loader = GraphLoader("/tmp/my_graph.json")
loader.save_graph(graph)

Summary

  • Grammar Initialization: codebase_rag/tools/language.py loads and caches Tree-Sitter language binaries to avoid repeated filesystem operations.
  • Language Parsers: Specialized modules under codebase_rag/parsers/ (Python, JS/TS, Rust) extract language-specific AST information using targeted Tree-Sitter queries.
  • Traversal Strategy: The PythonAstAnalyzerMixin and analogous classes implement _traverse_single_pass to ensure complete AST coverage via query-first or manual walk strategies.
  • Graph Construction: GraphUpdater.update_graph_from_ast() converts AST entities into typed nodes and relationships (CALLS, INHERITS, IMPORTS, DEFINE).
  • Persistence: GraphLoader serializes graphs to JSON and provides lookup utilities for downstream RAG consumption.

Frequently Asked Questions

What parsing technology does code-graph-rag use for multi-language support?

code-graph-rag uses Tree-Sitter as its foundation for multi-language parsing. According to the source code in codebase_rag/tools/language.py, the system discovers and loads compiled Tree-Sitter language grammars (like tree_sitter_python and tree_sitter_javascript) to create reusable Language objects that power the AST parsers.

How does the system handle different AST structures across programming languages?

The architecture delegates parsing to language-specific modules under codebase_rag/parsers/. Each language implements its own analyzer (such as PythonAstAnalyzerMixin for Python or type_inference.py for Rust) that understands the specific AST node types of that language. These analyzers use cached Tree-Sitter queries to extract relevant syntactic entities while maintaining a consistent interface for the graph construction phase.

What types of relationships are captured in the generated code graph?

The GraphUpdater class creates several relationship types including CALLS (function invocations), INHERITS (class inheritance), IMPORTS (module dependencies), and DEFINE (module-to-entity declarations). These relationships are stored as GraphRelRecord objects alongside GraphNodeRecord entities to create a fully navigable graph structure suitable for dependency analysis and RAG queries.

Can the generated graph be saved and reused across sessions?

Yes, the GraphLoader class in codebase_rag/graph_loader.py provides serialization capabilities. After construction via GraphUpdater, graphs can be persisted to JSON format using loader.save_graph(), then reloaded into memory later using load_graph(). The loader includes query methods like find_nodes_by_label to support efficient graph traversal during RAG operations.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →