# How Graphify’s Tree-sitter AST Extraction Pipeline Works Internally

> Explore Graphify's seven-stage Tree-sitter pipeline that parses files traverses ASTs and persists normalized nodes and edges to a graph store for deterministic knowledge graphs.

- Repository: [Graphify Labs/graphify](https://github.com/Graphify-Labs/graphify)
- Tags: internals
- Published: 2026-07-15

---

**Graphify converts source code into deterministic knowledge graphs using a seven-stage Tree-sitter pipeline that parses files, traverses ASTs, and persists normalized nodes and edges to a graph store.**

Graphify is an open-source knowledge graph builder that transforms raw source code into queryable graph structures. Its **Tree-sitter AST extraction pipeline** serves as the core engine, converting files from 29 supported programming languages into deterministic abstract syntax trees before normalizing them into nodes and edges. This architecture ensures that identical source files always produce identical graph fragments, enabling incremental updates and efficient caching across multiple extraction runs.

## The Seven-Stage Extraction Pipeline

The pipeline processes every source file through discrete stages, each handled by a dedicated module in the `graphify/src/` directory.

### Stage 1: Language Detection

The system maps file extensions to Tree-sitter language modules using the registry in [`graphify/src/languages/language_map.py`](https://github.com/Graphify-Labs/graphify/blob/main/graphify/src/languages/language_map.py). When encountering a `.js`, `.py`, or `.ts` file, the router identifies the appropriate grammar to load.

### Stage 2: Parser Loading

For the detected language, Graphify loads the corresponding compiled Tree-sitter grammar (e.g., `tree_sitter_javascript`, `tree_sitter_python`) via [`graphify/src/parsers/tree_sitter_loader.py`](https://github.com/Graphify-Labs/graphify/blob/main/graphify/src/parsers/tree_sitter_loader.py). This module manages the dynamic loading of the 29 language bindings listed in `uv.lock`.

### Stage 3: Source-File Parsing

The raw file content feeds into the Tree-sitter parser, producing a syntax tree object (`Tree`). This logic resides in [`graphify/src/parsers/tree_sitter_parser.py`](https://github.com/Graphify-Labs/graphify/blob/main/graphify/src/parsers/tree_sitter_parser.py), which wraps the native Tree-sitter `Parser` class and handles encoding detection.

### Stage 4: AST Traversal and Normalisation

The `ASTTraverser` class in [`graphify/src/ast/ast_traverser.py`](https://github.com/Graphify-Labs/graphify/blob/main/graphify/src/ast/ast_traverser.py) executes a depth-first walk of the Tree-sitter tree. It creates a Graphify **Node** for every Tree-sitter *named* node (functions, classes, statements), generating deterministic **Node IDs** by hashing the node's start byte, end byte, and type. The traverser also creates edges representing parent-child relationships from the Tree-sitter hierarchy.

### Stage 5: Edge Enrichment

After extracting the raw tree structure, [`graphify/src/ast/edge_enricher.py`](https://github.com/Graphify-Labs/graphify/blob/main/graphify/src/ast/edge_enricher.py) augments the graph with semantic edge stubs. These placeholders for *calls*, *references*, and *inherits* relationships are later populated by language-specific analyzers during the analysis phase.

### Stage 6: Persistence

The constructed nodes and edges are written to the backing store through [`graphify/src/storage/graph_store.py`](https://github.com/Graphify-Labs/graphify/blob/main/graphify/src/storage/graph_store.py). Graphify supports both **Neo4j** and **SQLite** backends, batch-inserting the AST fragments to enable fast querying for downstream analyses.

### Stage 7: Caching

To avoid re-parsing unchanged files, the pipeline checks [`graphify/src/cache/ast_cache.py`](https://github.com/Graphify-Labs/graphify/blob/main/graphify/src/cache/ast_cache.py) for existing ASTs. Parsed trees are serialized to disk in `/tmp/graphify/cache/`, keyed by content hash, allowing subsequent runs to skip redundant parsing work.

## Integration Architecture and Entry Points

Understanding how these stages connect reveals the pipeline's orchestration layer.

### Entry Point and File Routing

The process begins in [`graphify/main.py`](https://github.com/Graphify-Labs/graphify/blob/main/graphify/main.py), which iterates over the repository files and dispatches each to the appropriate handler. The `FileRouter` class in [`graphify/src/router/file_router.py`](https://github.com/Graphify-Labs/graphify/blob/main/graphify/src/router/file_router.py) determines whether a file requires Tree-sitter parsing (for code) or LLM-based semantic extraction (for documentation and images).

### Core Parser Classes

The **TreeSitterParser** class (in [`tree_sitter_parser.py`](https://github.com/Graphify-Labs/graphify/blob/main/tree_sitter_parser.py)) creates the initial `Tree` object, then hands it to the **ASTTraverser** (in [`ast_traverser.py`](https://github.com/Graphify-Labs/graphify/blob/main/ast_traverser.py)). The traverser emits Graphify-specific objects: `ASTNode` for vertices and `ASTEdge` for relationships, creating a clean abstraction layer over the raw Tree-sitter output.

### Deterministic Identity Generation

Node IDs are generated in [`graphify/src/utils/hash_utils.py`](https://github.com/Graphify-Labs/graphify/blob/main/graphify/src/utils/hash_utils.py) using deterministic hashing of node attributes. This guarantees that the same source file always yields the same graph fragment, allowing Graphify to merge multiple extraction runs without duplication or identity conflicts.

## Why Tree-sitter Powers the Pipeline

Graphify selected Tree-sitter as its parsing foundation for three critical reasons:

- **Speed:** Tree-sitter parses locally in less than 10 milliseconds for most source files, keeping extraction costs near zero.
- **Language Coverage:** The `uv.lock` file defines **29 supported languages**, from JavaScript to Rust, each loaded as a separate compiled grammar.
- **Deterministic Structure:** Unlike regular expression-based tokenizers, Tree-sitter guarantees a full, correct syntactic tree that Graphify can translate directly into graph edges without heuristic parsing.

## Practical Usage Examples

Below are concrete implementations showing how to interact with the pipeline programmatically.

This Python example demonstrates the high-level router interface:

```python

# example.py – parsing a JavaScript file and inserting its AST into the graph store

from graphify.src.router.file_router import FileRouter
from graphify.src.storage.graph_store import GraphStore

router = FileRouter()
store = GraphStore()

js_path = "src/app/main.js"
if router.is_code(js_path):
    ast = router.parse_with_tree_sitter(js_path)   # returns Graphify AST object

    store.save_ast(ast)                            # persists nodes & edges

```

For lower-level control, you can interact directly with the parser and traverser:

```typescript
// ts_example.ts – manual low‑level usage of the Tree‑sitter parser
import { TreeSitterParser } from "graphify/src/parsers/tree_sitter_parser";
import { ASTTraverser } from "graphify/src/ast/ast_traverser";

const source = await Deno.readTextFile("src/utils/helpers.ts");
const parser = new TreeSitterParser("typescript");
const tree = parser.parse(source);
const traverser = new ASTTraverser();
const graphAst = traverser.traverse(tree);
console.log(graphAst.nodes.length, "nodes extracted");

```

## Summary

- Graphify's pipeline consists of **seven stages**: language detection, parser loading, source-file parsing, AST traversal, edge enrichment, persistence, and caching.
- The **TreeSitterParser** and **ASTTraverser** classes in `graphify/src/parsers/` and `graphify/src/ast/` form the core extraction engine.
- **Deterministic hashing** in [`graphify/src/utils/hash_utils.py`](https://github.com/Graphify-Labs/graphify/blob/main/graphify/src/utils/hash_utils.py) ensures identical files produce identical graph fragments.
- Parsed ASTs are cached in `/tmp/graphify/cache/` via [`graphify/src/cache/ast_cache.py`](https://github.com/Graphify-Labs/graphify/blob/main/graphify/src/cache/ast_cache.py) to eliminate redundant processing.
- The system supports **29 languages** defined in `uv.lock` and persists results to Neo4j or SQLite via [`graphify/src/storage/graph_store.py`](https://github.com/Graphify-Labs/graphify/blob/main/graphify/src/storage/graph_store.py).

## Frequently Asked Questions

### How does Graphify ensure deterministic AST extraction across multiple runs?

Graphify generates node IDs using deterministic hashes of each node's start byte, end byte, and type in [`graphify/src/utils/hash_utils.py`](https://github.com/Graphify-Labs/graphify/blob/main/graphify/src/utils/hash_utils.py). This guarantees that identical source files always produce identical graph fragments, enabling safe merging of multiple extraction runs without duplication.

### What happens when Graphify encounters a file type it doesn't support?

The `FileRouter` class in [`graphify/src/router/file_router.py`](https://github.com/Graphify-Labs/graphify/blob/main/graphify/src/router/file_router.py) routes unsupported code files and non-code assets (like documentation or images) to an LLM-based semantic extractor rather than the Tree-sitter pipeline. This ensures comprehensive coverage even for languages not in the 29 supported Tree-sitter grammars listed in `uv.lock`.

### Where does Graphify store cached ASTs to avoid re-parsing?

The system serializes parsed ASTs to disk in `/tmp/graphify/cache/` using the `ASTCache` module in [`graphify/src/cache/ast_cache.py`](https://github.com/Graphify-Labs/graphify/blob/main/graphify/src/cache/ast_cache.py). Each cache entry is keyed by content hash, allowing subsequent runs to skip parsing for any file that hasn't changed since the last extraction.

### Which storage backends does Graphify support for the extracted graph data?

According to the source code in [`graphify/src/storage/graph_store.py`](https://github.com/Graphify-Labs/graphify/blob/main/graphify/src/storage/graph_store.py), Graphify persists nodes and edges to either **Neo4j** (for production graph databases) or **SQLite** (for local development). This dual-backend approach provides flexibility for different deployment scenarios.