How Code-Graph-RAG's Multi-Language Parser Utilizes Tree-Sitter for Universal Code Analysis
Code-Graph-RAG leverages Tree-sitter as its universal parsing engine through a language-agnostic pipeline that maps file extensions to specialized LanguageSpec objects, executes Tree-sitter queries against ASTs, and caches parsers for performance across Python, JavaScript, Rust, and other supported languages.
The open-source repository vitali87/code-graph-rag implements a sophisticated multi-language parsing system built entirely on Tree-sitter's incremental parsing library. By abstracting language-specific details into configurable specification objects and utilizing Tree-sitter's query syntax for node extraction, the system provides uniform code analysis capabilities across diverse programming languages without requiring separate parser implementations.
Language Detection and Specification Mapping
The parsing pipeline begins in codebase_rag/language_spec.py, where the system maps file extensions to LanguageSpec instances via get_language_spec() and get_language_for_extension(). Each LanguageSpec defines the complete parsing configuration for a specific language, including the Tree-sitter grammar package, query strings for entity extraction, and language-specific name resolution helpers.
For every supported language—ranging from Python and JavaScript to Rust, Java, C/C++, PHP, and Go—the specification contains:
- Grammar references: Lazy-loaded Tree-sitter language objects (e.g.,
tree_sitter_python,tree_sitter_rust) - Query definitions: Tree-sitter query strings stored in
function_query,class_query, andcall_queryfields - Name extractors: Language-specific functions like
_python_get_name,_js_get_name, and_rust_get_namethat navigate AST node fields
Grammar Loading and Lazy Parser Initialization
Tree-sitter grammars are imported lazily to minimize memory overhead and startup time. When get_language_spec() resolves a file extension to a specific language, the system loads the appropriate Tree-sitter grammar package and constructs a tree_sitter.Language object, which is then cached for subsequent parsing operations.
This lazy-loading strategy ensures that parsers are instantiated only when needed. The test suite in codebase_rag/tests/test_lazy_parser_loading.py verifies that grammar compilation occurs on-demand rather than during system initialization, significantly improving performance when processing large repositories containing multiple languages.
AST Construction and Query-Driven Extraction
Once a language is identified, the system creates a tree_sitter.Parser instance configured with the appropriate Language object. The parser processes source files via parser.parse(source_bytes), generating a Tree-sitter syntax tree where the root node serves as the entry point for all analyses.
Entity extraction relies on Tree-sitter queries compiled from the LanguageSpec definitions. The system creates tree_sitter.Query objects from the specification's query strings and executes them against the AST using a QueryCursor. These queries precisely target nodes representing:
- Function definitions
- Class declarations
- Import statements
- Call expressions
For example, the function_query string defined in language_spec.py (around lines 398-424 for Rust specifications) captures function declaration nodes while binding identifier names to capture groups.
Name Resolution and Symbol Extraction
After Tree-sitter queries identify relevant AST nodes, language-specific helper functions extract identifier text by navigating node fields. The system implements specialized extractors such as _python_get_name, _js_get_name, and _rust_get_name, each utilizing child_by_field_name("name") to retrieve identifier nodes from their parent declarations.
When language-specific handlers aren't available, the fallback _generic_get_name function scans common field name patterns across languages. This dual approach ensures robust symbol extraction regardless of whether the target language uses conventional AST structures or language-specific naming conventions.
Additionally, *_file_to_module helpers transform file paths into dotted module names according to language conventions—stripping __init__.py for Python, index.js for JavaScript, or mod.rs for Rust—to maintain consistent module resolution across the codebase.
Pipeline Integration via ProcessorFactory
The ProcessorFactory class in codebase_rag/parsers/factory.py orchestrates the complete analysis pipeline by instantiating specialized processors that utilize the Tree-sitter infrastructure:
- ImportProcessor: Discovers cross-file dependencies using Tree-sitter import queries
- StructureProcessor: Maps class hierarchies and module organization
- DefinitionProcessor: Indexes function and class definitions with precise source locations
- CallProcessor: Builds the call graph by analyzing invocation sites via
call_queryexecution
Each processor receives the LanguageSpec and compiled Tree-sitter queries, enabling uniform traversal of ASTs while respecting language-specific semantics. This architecture allows Code-Graph-RAG to construct comprehensive call graphs spanning multiple programming languages within a single repository.
Practical Implementation Examples
The following example demonstrates parsing a Python file and extracting function names using the Tree-sitter infrastructure:
from pathlib import Path
from codebase_rag.language_spec import get_language_spec, LanguageSpec
from tree_sitter import Language, Parser, Query
def extract_python_functions(file_path: Path):
# Resolve the language spec based on file extension
spec: LanguageSpec | None = get_language_spec('.py')
assert spec, "Unsupported file type"
# Load the Tree-sitter language (cached internally)
from tree_sitter_python import language as py_lang
language = Language(py_lang.language())
# Initialize parser and parse source bytes
parser = Parser()
parser.set_language(language)
source = file_path.read_bytes()
tree = parser.parse(source)
# Compile function query from specification
query = Query(language, spec.function_query)
# Execute query and extract function names
cursor = query.exec(tree.root_node)
functions = []
for match in cursor:
name_node = match.captures[0][0] # First capture contains identifier
functions.append(name_node.text.decode('utf8'))
return functions
For high-level pipeline usage, the ProcessorFactory provides automatic Tree-sitter integration:
from codebase_rag.parsers.factory import ProcessorFactory
factory = ProcessorFactory(
ingestor=my_ingestor,
repo_path=Path('/my/repo'),
project_name='myproject',
queries=language_queries, # Mapping[SupportedLanguage, LanguageQueries]
function_registry=my_registry,
simple_name_lookup=my_lookup,
ast_cache=my_ast_cache,
)
# Factory lazily creates CallProcessor with Tree-sitter configuration
call_processor = factory.call_processor
call_graph = call_processor.build_call_graph()
Summary
- Tree-sitter serves as the universal front-end for Code-Graph-RAG's multi-language parsing, converting source files into language-agnostic ASTs through the
tree_sitter.Parserinterface. - Language specifications centralize parsing logic in
codebase_rag/language_spec.py, defining queries, grammars, and name extraction helpers for each supported programming language. - Query-driven extraction utilizes compiled Tree-sitter queries (
function_query,class_query,call_query) to identify code entities without manual AST traversal logic. - Lazy loading and caching optimize performance by instantiating Tree-sitter parsers and compiling queries only when specific languages are encountered during repository analysis.
- ProcessorFactory orchestrates cross-language analysis by providing Tree-sitter-enabled processors that build import graphs, definition indexes, and call graphs uniformly across Python, JavaScript, Rust, and other languages.
Frequently Asked Questions
What is Tree-sitter and why does Code-Graph-RAG use it?
Tree-sitter is an incremental parsing library that generates syntax trees for source code, enabling fast re-parsing as files change. Code-Graph-RAG utilizes Tree-sitter because it provides robust, error-tolerant parsing across multiple programming languages through a consistent C API and Python bindings, eliminating the need to maintain separate parsers for each supported language.
How does the parser handle language-specific syntax differences?
The system abstracts language differences through LanguageSpec objects defined in codebase_rag/language_spec.py. Each specification contains Tree-sitter query strings tailored to that language's grammar, along with specialized name extraction functions (like _rust_get_name or _js_get_name) that understand language-specific AST node structures and field names.
What are Tree-sitter queries and how do they extract code entities?
Tree-sitter queries are pattern-matching expressions written in a Lisp-like syntax that select specific nodes from the AST. In Code-Graph-RAG, queries stored in function_query, class_query, and call_query fields define patterns like (function_declaration name: (identifier) @name) to capture function names. The system compiles these strings into tree_sitter.Query objects and executes them via QueryCursor to extract identifiers without manual tree traversal.
How does the system optimize performance when parsing large codebases?
Code-Graph-RAG implements lazy loading of Tree-sitter grammars and caches parser instances per language to avoid repeated compilation overhead. The ProcessorFactory reuses compiled queries and AST caches across multiple files, while the architecture ensures that Tree-sitter Language objects and Parser instances are constructed only when processing files of that specific type, as verified by codebase_rag/tests/test_lazy_parser_loading.py.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →