# How Tree-sitter Parser Extracts Code Relationships into the Memgraph Knowledge Graph

> Discover how Tree-sitter parser extracts code relationships into Memgraph knowledge graph. Learn about AST queries, node batching, and efficient data ingestion for code analysis.

- Repository: [Vitali Avagyan/code-graph-rag](https://github.com/vitali87/code-graph-rag)
- Tags: how-to-guide
- Published: 2026-08-20

---

**The pipeline initializes Tree-sitter parsers per language, runs syntactic queries against ASTs to identify definitions and calls, then batches nodes and relationships to Memgraph via the FilteringIngestor.**

The `vitali87/code-graph-rag` repository implements a multi-pass ingestion system that transforms source code into a queryable knowledge graph. The **Tree-sitter parser** serves as the primary syntactic engine, supplemented by hybrid front-ends for languages requiring additional semantic analysis. This article breaks down exactly how code relationships flow from raw text into a structured Memgraph database.

---

## Overview of the Ingestion Pipeline

The ingestion process follows a **three-pass architecture** orchestrated by the `GraphUpdater` class. Tree-sitter processing occurs in **Pass 2**, where source files are parsed, queried, and transformed into graph elements.

| Pass | Purpose | Key Activity |
|------|---------|------------|
| Pass 1 | File discovery | Identify source files by extension |
| **Pass 2** | **AST analysis** | **Parse with Tree-sitter, run queries, extract spans** |
| Pass 3 | Relationship resolution | Link call sites to definitions, resolve macros |

The `GraphUpdater.run()` method drives this flow, with `_process_files()` handling the Tree-sitter-specific work.

---

## Step 1: Parser Initialization and Language Loading

At startup, [`parser_loader.py`](https://github.com/vitali87/code-graph-rag/blob/main/parser_loader.py) constructs a **dedicated `Parser` object** for each supported language. These parsers wrap compiled Tree-sitter grammar files (`.so` shared libraries).

```python

# From codebase_rag/parsers/parser_loader.py

from tree_sitter import Language, Parser

# Example: loading Python grammar

python_lang = Language('build/grammars.so', 'python')
parser = Parser()
parser.set_language(python_lang)

```

The `COMBINED_FUNC_CLASS_IMPORT_QUERIES` dictionary stores **Tree-sitter query strings** for each language. These queries match syntactic patterns for:

- Function and method definitions
- Class declarations
- Import statements
- Call expressions
- Inheritance and implementation clauses

---

## Step 2: Parsing Source Files into ASTs

In [`codebase_rag/graph_updater.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/graph_updater.py), the `_process_files()` method iterates over source files and invokes `Parser.parse()`:

```python

# Simplified excerpt showing core parsing flow

source_bytes = Path('example.py').read_bytes()
tree = parser.parse(source_bytes)  # Returns tree_sitter.Tree

root_node = tree.root_node         # tree_sitter.Node

```

The parser returns a **`Tree`** object containing the complete **AST (Abstract Syntax Tree)**. Each node in this tree knows its exact byte ranges and line/column positions, enabling precise span extraction for later relationship linking.

---

## Step 3: Executing Tree-sitter Queries

A `QueryCursor` executes language-specific queries against the AST. The cursor yields **matches** containing captured nodes with their positions:

```python

# Example query to find function definitions

query = parser.language.query(
    "(function_definition name: (identifier) @func_name)"
)
cursor = query.exec(tree.root_node)

for match in cursor:
    # Each match contains the captured node with exact span

    func_name_node = match.captures[0][0]
    qualified_name = func_name_node.text.decode()
    start_line = func_name_node.start_point[0] + 1
    end_line = func_name_node.end_point[0] + 1

```

The `ProcessorFactory.structure_processor.identify_structure()` method coordinates this cursor iteration, delegating to specialized processors for definitions, imports, and calls.

---

## Step 4: Registering Nodes in the Graph

For each captured definition, the system creates a **typed node** using `FilteringIngestor.ensure_node_batch()`. Node labels are drawn from [`constants/graph.py`](https://github.com/vitali87/code-graph-rag/blob/main/constants/graph.py):

- `NodeLabel.FUNCTION`
- `NodeLabel.METHOD`
- `NodeLabel.CLASS`
- `NodeLabel.MODULE`

```python

# Node creation pattern from GraphUpdater Pass 2 processing

self._sink.ensure_node_batch(
    cs.NodeLabel.FUNCTION,
    {
        cs.KEY_QUALIFIED_NAME: qualified_name,
        cs.KEY_START_LINE: start_line,
        cs.KEY_END_LINE: end_line
    }
)

```

The `_sink` reference points to the `FilteringIngestor`, which wraps the raw Memgraph connection and respects capture filters before batching operations.

---

## Step 5: Emitting Code Relationships

When queries identify **call sites, imports, inheritance, or implementations**, the processor emits **typed relationships**:

```python

# Relationship emission pattern (from _resolve_hybrid_macro_calls, lines 66-77)

self._sink.ensure_relationship_batch(
    (cs.NodeLabel.FUNCTION, cs.KEY_QUALIFIED_NAME, caller_qn),
    cs.RelationshipType.CALLS,
    (cs.NodeLabel.FUNCTION, cs.KEY_QUALIFIED_NAME, callee_qn)
)

```

Key relationship types from [`constants.py`](https://github.com/vitali87/code-graph-rag/blob/main/constants.py) include:

- `RelationshipType.CALLS` — function/method invocations
- `RelationshipType.IMPORTS` — module dependencies
- `RelationshipType.INHERITS` — class inheritance
- `RelationshipType.IMPLEMENTS` — interface implementations

---

## Step 6: Computing Tightest-Containing Spans

Tree-sitter's position data enables **precise span computation** for relationship targeting. The `_tightest_containing_span` method (lines 51-64 in [`graph_updater.py`](https://github.com/vitali87/code-graph-rag/blob/main/graph_updater.py)) calculates the minimal source range containing a call expression:

```python
def _tightest_containing_span(self, call_node, file_path):
    """
    Uses Tree-sitter node coordinates to find the exact
    line range for a call site, enabling accurate
    relationship linking to definitions.
    """
    # Tree-sitter provides byte offsets and line/column pairs

    start_line = call_node.start_point[0] + 1
    end_line = call_node.end_point[0] + 1
    return (file_path, start_line, end_line)

```

This precision is critical for **macro expansion resolution** and **hybrid analysis merging**.

---

## Step 7: Flushing Data to Memgraph

After all passes complete, `GraphUpdater.run()` calls `self.ingestor.flush_all()` (lines 77-78):

```python

# End of ingestion pipeline

def run(self):
    # ... Pass 1, 2, 3 processing ...

    self.ingestor.flush_all()  # Sends batched Cypher to Memgraph

```

The `FilteringIngestor` (from [`services/resource_cleanup.py`](https://github.com/vitali87/code-graph-rag/blob/main/services/resource_cleanup.py)) batches nodes and relationships into **Cypher commands**, then executes them against the Memgraph Bolt protocol.

---

## Hybrid Front-End Integration

Tree-sitter handles **pure syntactic analysis**. For languages requiring semantic information, **hybrid front-ends** supplement the pipeline:

| Language | Front-End | Purpose |
|----------|-----------|---------|
| C/C++ | libclang | Macro expansion resolution |
| C# | Roslyn | Type information |

| Go | `go/types` | Package-level analysis |

The hybrid data joins Tree-sitter spans in `_resolve_hybrid_macro_calls()` and `_resolve_hybrid_expansion_calls()`, enriching the graph with information Tree-sitter alone cannot provide.

---

## Capture Filtering for Selective Ingestion

The `FilteringIngestor` supports **relationship-type filtering** via `self.capture`. Before any relationship reaches Memgraph, the filter checks whether that relationship type is enabled:

```python

# Conceptual filtering check

if relationship_type in self.enabled_captures:
    self._inner_ingestor.ensure_relationship_batch(source, rel_type, target)

```

This allows users to ingest only call graphs, exclude imports, or focus on inheritance hierarchies without modifying parser code.

---

## Summary

- **Tree-sitter parsers** are instantiated per-language in [`parser_loader.py`](https://github.com/vitali87/code-graph-rag/blob/main/parser_loader.py), loading compiled grammars and query collections.

- **AST parsing** occurs in `GraphUpdater._process_files()` during Pass 2, producing `tree_sitter.Node` trees with exact position data.

- **Query cursors** execute patterns from `COMBINED_FUNC_CLASS_IMPORT_QUERIES`, extracting definitions, calls, imports, and inheritance.

- **Nodes and relationships** are batched via `FilteringIngestor.ensure_node_batch()` and `ensure_relationship_batch()`, with types from [`constants.py`](https://github.com/vitali87/code-graph-rag/blob/main/constants.py).

- **Span precision** from Tree-sitter enables accurate relationship targeting and hybrid data merging for macro/type resolution.

- **Final flush** via `ingestor.flush_all()` commits all batched Cypher commands to Memgraph.

---

## Frequently Asked Questions

### What file formats does the Tree-sitter parser support?

The parser supports any language with a compiled Tree-sitter grammar in the `grammars/` directory, including Python, JavaScript, TypeScript, Ruby, Go, Rust, C, C++, C#, and Java. Extensions are mapped to languages in [`codebase_rag/constants/languages.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/constants/languages.py). Unsupported extensions fall back to basic file-level node creation without AST analysis.

### How does the system handle macro calls in C/C++?

C and C++ use a **hybrid front-end** combining Tree-sitter with libclang. Tree-sitter identifies the syntactic span of macro invocations, while libclang resolves the actual expansion. The `_resolve_hybrid_macro_calls()` method joins this information, emitting `CALLS` relationships from the expansion site to the final callee with accurate line ranges.

### Can I disable specific relationship types during ingestion?

Yes. The `FilteringIngestor` respects capture filters set via configuration. Disable unwanted relationship types—such as `IMPORTS` or `INHERITS`—before running `GraphUpdater`, and those edges will be excluded from the batch sent to Memgraph without requiring parser modifications.

### Where is the relationship type `CALLS` defined?

All relationship types are enumerated in [`codebase_rag/constants/graph.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/constants/graph.py) as the `RelationshipType` class. Available types include `CALLS`, `IMPORTS`, `INHERITS`, `IMPLEMENTS`, `CONTAINS`, and `REFERENCES`. Node labels like `FUNCTION` and `CLASS` are defined in the same file under `NodeLabel`.