# How Code-Graph-RAG's Multi-Language Parser Utilizes Tree-Sitter for Universal Code Analysis

> Discover how Code-Graph-RAG's multi-language parser uses Tree-sitter for universal code analysis. Learn about its efficient AST querying and caching across Python, JavaScript, Rust, and more.

- Repository: [Vitali Avagyan/code-graph-rag](https://github.com/vitali87/code-graph-rag)
- Tags: internals
- Published: 2026-08-18

---

**Code-Graph-RAG leverages Tree-sitter as its universal parsing engine through a language-agnostic pipeline that maps file extensions to specialized `LanguageSpec` objects, executes Tree-sitter queries against ASTs, and caches parsers for performance across Python, JavaScript, Rust, and other supported languages.**

The open-source repository `vitali87/code-graph-rag` implements a sophisticated multi-language parsing system built entirely on Tree-sitter's incremental parsing library. By abstracting language-specific details into configurable specification objects and utilizing Tree-sitter's query syntax for node extraction, the system provides uniform code analysis capabilities across diverse programming languages without requiring separate parser implementations.

## Language Detection and Specification Mapping

The parsing pipeline begins in [`codebase_rag/language_spec.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/language_spec.py), where the system maps file extensions to `LanguageSpec` instances via `get_language_spec()` and `get_language_for_extension()`. Each `LanguageSpec` defines the complete parsing configuration for a specific language, including the Tree-sitter grammar package, query strings for entity extraction, and language-specific name resolution helpers.

For every supported language—ranging from Python and JavaScript to Rust, Java, C/C++, PHP, and Go—the specification contains:

- **Grammar references**: Lazy-loaded Tree-sitter language objects (e.g., `tree_sitter_python`, `tree_sitter_rust`)
- **Query definitions**: Tree-sitter query strings stored in `function_query`, `class_query`, and `call_query` fields
- **Name extractors**: Language-specific functions like `_python_get_name`, `_js_get_name`, and `_rust_get_name` that navigate AST node fields

## Grammar Loading and Lazy Parser Initialization

Tree-sitter grammars are imported lazily to minimize memory overhead and startup time. When `get_language_spec()` resolves a file extension to a specific language, the system loads the appropriate Tree-sitter grammar package and constructs a `tree_sitter.Language` object, which is then cached for subsequent parsing operations.

This lazy-loading strategy ensures that parsers are instantiated only when needed. The test suite in [`codebase_rag/tests/test_lazy_parser_loading.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/tests/test_lazy_parser_loading.py) verifies that grammar compilation occurs on-demand rather than during system initialization, significantly improving performance when processing large repositories containing multiple languages.

## AST Construction and Query-Driven Extraction

Once a language is identified, the system creates a `tree_sitter.Parser` instance configured with the appropriate `Language` object. The parser processes source files via `parser.parse(source_bytes)`, generating a Tree-sitter syntax tree where the root node serves as the entry point for all analyses.

Entity extraction relies on **Tree-sitter queries** compiled from the `LanguageSpec` definitions. The system creates `tree_sitter.Query` objects from the specification's query strings and executes them against the AST using a `QueryCursor`. These queries precisely target nodes representing:

- Function definitions
- Class declarations
- Import statements
- Call expressions

For example, the `function_query` string defined in [`language_spec.py`](https://github.com/vitali87/code-graph-rag/blob/main/language_spec.py) (around lines 398-424 for Rust specifications) captures function declaration nodes while binding identifier names to capture groups.

## Name Resolution and Symbol Extraction

After Tree-sitter queries identify relevant AST nodes, language-specific helper functions extract identifier text by navigating node fields. The system implements specialized extractors such as `_python_get_name`, `_js_get_name`, and `_rust_get_name`, each utilizing `child_by_field_name("name")` to retrieve identifier nodes from their parent declarations.

When language-specific handlers aren't available, the fallback `_generic_get_name` function scans common field name patterns across languages. This dual approach ensures robust symbol extraction regardless of whether the target language uses conventional AST structures or language-specific naming conventions.

Additionally, `*_file_to_module` helpers transform file paths into dotted module names according to language conventions—stripping [`__init__.py`](https://github.com/vitali87/code-graph-rag/blob/main/__init__.py) for Python, [`index.js`](https://github.com/vitali87/code-graph-rag/blob/main/index.js) for JavaScript, or [`mod.rs`](https://github.com/vitali87/code-graph-rag/blob/main/mod.rs) for Rust—to maintain consistent module resolution across the codebase.

## Pipeline Integration via ProcessorFactory

The `ProcessorFactory` class in [`codebase_rag/parsers/factory.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/parsers/factory.py) orchestrates the complete analysis pipeline by instantiating specialized processors that utilize the Tree-sitter infrastructure:

- **ImportProcessor**: Discovers cross-file dependencies using Tree-sitter import queries
- **StructureProcessor**: Maps class hierarchies and module organization
- **DefinitionProcessor**: Indexes function and class definitions with precise source locations
- **CallProcessor**: Builds the call graph by analyzing invocation sites via `call_query` execution

Each processor receives the `LanguageSpec` and compiled Tree-sitter queries, enabling uniform traversal of ASTs while respecting language-specific semantics. This architecture allows Code-Graph-RAG to construct comprehensive call graphs spanning multiple programming languages within a single repository.

## Practical Implementation Examples

The following example demonstrates parsing a Python file and extracting function names using the Tree-sitter infrastructure:

```python
from pathlib import Path
from codebase_rag.language_spec import get_language_spec, LanguageSpec
from tree_sitter import Language, Parser, Query

def extract_python_functions(file_path: Path):
    # Resolve the language spec based on file extension

    spec: LanguageSpec | None = get_language_spec('.py')
    assert spec, "Unsupported file type"

    # Load the Tree-sitter language (cached internally)

    from tree_sitter_python import language as py_lang
    language = Language(py_lang.language())

    # Initialize parser and parse source bytes

    parser = Parser()
    parser.set_language(language)
    source = file_path.read_bytes()
    tree = parser.parse(source)

    # Compile function query from specification

    query = Query(language, spec.function_query)

    # Execute query and extract function names

    cursor = query.exec(tree.root_node)
    functions = []
    for match in cursor:
        name_node = match.captures[0][0]  # First capture contains identifier

        functions.append(name_node.text.decode('utf8'))

    return functions

```

For high-level pipeline usage, the `ProcessorFactory` provides automatic Tree-sitter integration:

```python
from codebase_rag.parsers.factory import ProcessorFactory

factory = ProcessorFactory(
    ingestor=my_ingestor,
    repo_path=Path('/my/repo'),
    project_name='myproject',
    queries=language_queries,           # Mapping[SupportedLanguage, LanguageQueries]

    function_registry=my_registry,
    simple_name_lookup=my_lookup,
    ast_cache=my_ast_cache,
)

# Factory lazily creates CallProcessor with Tree-sitter configuration

call_processor = factory.call_processor
call_graph = call_processor.build_call_graph()

```

## Summary

- **Tree-sitter serves as the universal front-end** for Code-Graph-RAG's multi-language parsing, converting source files into language-agnostic ASTs through the `tree_sitter.Parser` interface.
- **Language specifications centralize parsing logic** in [`codebase_rag/language_spec.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/language_spec.py), defining queries, grammars, and name extraction helpers for each supported programming language.
- **Query-driven extraction** utilizes compiled Tree-sitter queries (`function_query`, `class_query`, `call_query`) to identify code entities without manual AST traversal logic.
- **Lazy loading and caching** optimize performance by instantiating Tree-sitter parsers and compiling queries only when specific languages are encountered during repository analysis.
- **ProcessorFactory orchestrates cross-language analysis** by providing Tree-sitter-enabled processors that build import graphs, definition indexes, and call graphs uniformly across Python, JavaScript, Rust, and other languages.

## Frequently Asked Questions

### What is Tree-sitter and why does Code-Graph-RAG use it?

Tree-sitter is an incremental parsing library that generates syntax trees for source code, enabling fast re-parsing as files change. Code-Graph-RAG utilizes Tree-sitter because it provides robust, error-tolerant parsing across multiple programming languages through a consistent C API and Python bindings, eliminating the need to maintain separate parsers for each supported language.

### How does the parser handle language-specific syntax differences?

The system abstracts language differences through `LanguageSpec` objects defined in [`codebase_rag/language_spec.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/language_spec.py). Each specification contains Tree-sitter query strings tailored to that language's grammar, along with specialized name extraction functions (like `_rust_get_name` or `_js_get_name`) that understand language-specific AST node structures and field names.

### What are Tree-sitter queries and how do they extract code entities?

Tree-sitter queries are pattern-matching expressions written in a Lisp-like syntax that select specific nodes from the AST. In Code-Graph-RAG, queries stored in `function_query`, `class_query`, and `call_query` fields define patterns like `(function_declaration name: (identifier) @name)` to capture function names. The system compiles these strings into `tree_sitter.Query` objects and executes them via `QueryCursor` to extract identifiers without manual tree traversal.

### How does the system optimize performance when parsing large codebases?

Code-Graph-RAG implements lazy loading of Tree-sitter grammars and caches parser instances per language to avoid repeated compilation overhead. The `ProcessorFactory` reuses compiled queries and AST caches across multiple files, while the architecture ensures that Tree-sitter `Language` objects and `Parser` instances are constructed only when processing files of that specific type, as verified by [`codebase_rag/tests/test_lazy_parser_loading.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/tests/test_lazy_parser_loading.py).