How Tree-sitter Parser Extracts Code Relationships into the Memgraph Knowledge Graph

The pipeline initializes Tree-sitter parsers per language, runs syntactic queries against ASTs to identify definitions and calls, then batches nodes and relationships to Memgraph via the FilteringIngestor.

The vitali87/code-graph-rag repository implements a multi-pass ingestion system that transforms source code into a queryable knowledge graph. The Tree-sitter parser serves as the primary syntactic engine, supplemented by hybrid front-ends for languages requiring additional semantic analysis. This article breaks down exactly how code relationships flow from raw text into a structured Memgraph database.


Overview of the Ingestion Pipeline

The ingestion process follows a three-pass architecture orchestrated by the GraphUpdater class. Tree-sitter processing occurs in Pass 2, where source files are parsed, queried, and transformed into graph elements.

Pass Purpose Key Activity
Pass 1 File discovery Identify source files by extension
Pass 2 AST analysis Parse with Tree-sitter, run queries, extract spans
Pass 3 Relationship resolution Link call sites to definitions, resolve macros

The GraphUpdater.run() method drives this flow, with _process_files() handling the Tree-sitter-specific work.


Step 1: Parser Initialization and Language Loading

At startup, parser_loader.py constructs a dedicated Parser object for each supported language. These parsers wrap compiled Tree-sitter grammar files (.so shared libraries).


# From codebase_rag/parsers/parser_loader.py

from tree_sitter import Language, Parser

# Example: loading Python grammar

python_lang = Language('build/grammars.so', 'python')
parser = Parser()
parser.set_language(python_lang)

The COMBINED_FUNC_CLASS_IMPORT_QUERIES dictionary stores Tree-sitter query strings for each language. These queries match syntactic patterns for:

  • Function and method definitions
  • Class declarations
  • Import statements
  • Call expressions
  • Inheritance and implementation clauses

Step 2: Parsing Source Files into ASTs

In codebase_rag/graph_updater.py, the _process_files() method iterates over source files and invokes Parser.parse():


# Simplified excerpt showing core parsing flow

source_bytes = Path('example.py').read_bytes()
tree = parser.parse(source_bytes)  # Returns tree_sitter.Tree

root_node = tree.root_node         # tree_sitter.Node

The parser returns a Tree object containing the complete AST (Abstract Syntax Tree). Each node in this tree knows its exact byte ranges and line/column positions, enabling precise span extraction for later relationship linking.


Step 3: Executing Tree-sitter Queries

A QueryCursor executes language-specific queries against the AST. The cursor yields matches containing captured nodes with their positions:


# Example query to find function definitions

query = parser.language.query(
    "(function_definition name: (identifier) @func_name)"
)
cursor = query.exec(tree.root_node)

for match in cursor:
    # Each match contains the captured node with exact span

    func_name_node = match.captures[0][0]
    qualified_name = func_name_node.text.decode()
    start_line = func_name_node.start_point[0] + 1
    end_line = func_name_node.end_point[0] + 1

The ProcessorFactory.structure_processor.identify_structure() method coordinates this cursor iteration, delegating to specialized processors for definitions, imports, and calls.


Step 4: Registering Nodes in the Graph

For each captured definition, the system creates a typed node using FilteringIngestor.ensure_node_batch(). Node labels are drawn from constants/graph.py:

  • NodeLabel.FUNCTION
  • NodeLabel.METHOD
  • NodeLabel.CLASS
  • NodeLabel.MODULE

# Node creation pattern from GraphUpdater Pass 2 processing

self._sink.ensure_node_batch(
    cs.NodeLabel.FUNCTION,
    {
        cs.KEY_QUALIFIED_NAME: qualified_name,
        cs.KEY_START_LINE: start_line,
        cs.KEY_END_LINE: end_line
    }
)

The _sink reference points to the FilteringIngestor, which wraps the raw Memgraph connection and respects capture filters before batching operations.


Step 5: Emitting Code Relationships

When queries identify call sites, imports, inheritance, or implementations, the processor emits typed relationships:


# Relationship emission pattern (from _resolve_hybrid_macro_calls, lines 66-77)

self._sink.ensure_relationship_batch(
    (cs.NodeLabel.FUNCTION, cs.KEY_QUALIFIED_NAME, caller_qn),
    cs.RelationshipType.CALLS,
    (cs.NodeLabel.FUNCTION, cs.KEY_QUALIFIED_NAME, callee_qn)
)

Key relationship types from constants.py include:

  • RelationshipType.CALLS — function/method invocations
  • RelationshipType.IMPORTS — module dependencies
  • RelationshipType.INHERITS — class inheritance
  • RelationshipType.IMPLEMENTS — interface implementations

Step 6: Computing Tightest-Containing Spans

Tree-sitter's position data enables precise span computation for relationship targeting. The _tightest_containing_span method (lines 51-64 in graph_updater.py) calculates the minimal source range containing a call expression:

def _tightest_containing_span(self, call_node, file_path):
    """
    Uses Tree-sitter node coordinates to find the exact
    line range for a call site, enabling accurate
    relationship linking to definitions.
    """
    # Tree-sitter provides byte offsets and line/column pairs

    start_line = call_node.start_point[0] + 1
    end_line = call_node.end_point[0] + 1
    return (file_path, start_line, end_line)

This precision is critical for macro expansion resolution and hybrid analysis merging.


Step 7: Flushing Data to Memgraph

After all passes complete, GraphUpdater.run() calls self.ingestor.flush_all() (lines 77-78):


# End of ingestion pipeline

def run(self):
    # ... Pass 1, 2, 3 processing ...

    self.ingestor.flush_all()  # Sends batched Cypher to Memgraph

The FilteringIngestor (from services/resource_cleanup.py) batches nodes and relationships into Cypher commands, then executes them against the Memgraph Bolt protocol.


Hybrid Front-End Integration

Tree-sitter handles pure syntactic analysis. For languages requiring semantic information, hybrid front-ends supplement the pipeline:

Language Front-End Purpose
C/C++ libclang Macro expansion resolution
C# Roslyn Type information

| Go | go/types | Package-level analysis |

The hybrid data joins Tree-sitter spans in _resolve_hybrid_macro_calls() and _resolve_hybrid_expansion_calls(), enriching the graph with information Tree-sitter alone cannot provide.


Capture Filtering for Selective Ingestion

The FilteringIngestor supports relationship-type filtering via self.capture. Before any relationship reaches Memgraph, the filter checks whether that relationship type is enabled:


# Conceptual filtering check

if relationship_type in self.enabled_captures:
    self._inner_ingestor.ensure_relationship_batch(source, rel_type, target)

This allows users to ingest only call graphs, exclude imports, or focus on inheritance hierarchies without modifying parser code.


Summary

  • Tree-sitter parsers are instantiated per-language in parser_loader.py, loading compiled grammars and query collections.

  • AST parsing occurs in GraphUpdater._process_files() during Pass 2, producing tree_sitter.Node trees with exact position data.

  • Query cursors execute patterns from COMBINED_FUNC_CLASS_IMPORT_QUERIES, extracting definitions, calls, imports, and inheritance.

  • Nodes and relationships are batched via FilteringIngestor.ensure_node_batch() and ensure_relationship_batch(), with types from constants.py.

  • Span precision from Tree-sitter enables accurate relationship targeting and hybrid data merging for macro/type resolution.

  • Final flush via ingestor.flush_all() commits all batched Cypher commands to Memgraph.


Frequently Asked Questions

What file formats does the Tree-sitter parser support?

The parser supports any language with a compiled Tree-sitter grammar in the grammars/ directory, including Python, JavaScript, TypeScript, Ruby, Go, Rust, C, C++, C#, and Java. Extensions are mapped to languages in codebase_rag/constants/languages.py. Unsupported extensions fall back to basic file-level node creation without AST analysis.

How does the system handle macro calls in C/C++?

C and C++ use a hybrid front-end combining Tree-sitter with libclang. Tree-sitter identifies the syntactic span of macro invocations, while libclang resolves the actual expansion. The _resolve_hybrid_macro_calls() method joins this information, emitting CALLS relationships from the expansion site to the final callee with accurate line ranges.

Can I disable specific relationship types during ingestion?

Yes. The FilteringIngestor respects capture filters set via configuration. Disable unwanted relationship types—such as IMPORTS or INHERITS—before running GraphUpdater, and those edges will be excluded from the batch sent to Memgraph without requiring parser modifications.

Where is the relationship type CALLS defined?

All relationship types are enumerated in codebase_rag/constants/graph.py as the RelationshipType class. Available types include CALLS, IMPORTS, INHERITS, IMPLEMENTS, CONTAINS, and REFERENCES. Node labels like FUNCTION and CLASS are defined in the same file under NodeLabel.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →