How Graphify’s Tree-sitter AST Extraction Pipeline Works Internally

Graphify converts source code into deterministic knowledge graphs using a seven-stage Tree-sitter pipeline that parses files, traverses ASTs, and persists normalized nodes and edges to a graph store.

Graphify is an open-source knowledge graph builder that transforms raw source code into queryable graph structures. Its Tree-sitter AST extraction pipeline serves as the core engine, converting files from 29 supported programming languages into deterministic abstract syntax trees before normalizing them into nodes and edges. This architecture ensures that identical source files always produce identical graph fragments, enabling incremental updates and efficient caching across multiple extraction runs.

The Seven-Stage Extraction Pipeline

The pipeline processes every source file through discrete stages, each handled by a dedicated module in the graphify/src/ directory.

Stage 1: Language Detection

The system maps file extensions to Tree-sitter language modules using the registry in graphify/src/languages/language_map.py. When encountering a .js, .py, or .ts file, the router identifies the appropriate grammar to load.

Stage 2: Parser Loading

For the detected language, Graphify loads the corresponding compiled Tree-sitter grammar (e.g., tree_sitter_javascript, tree_sitter_python) via graphify/src/parsers/tree_sitter_loader.py. This module manages the dynamic loading of the 29 language bindings listed in uv.lock.

Stage 3: Source-File Parsing

The raw file content feeds into the Tree-sitter parser, producing a syntax tree object (Tree). This logic resides in graphify/src/parsers/tree_sitter_parser.py, which wraps the native Tree-sitter Parser class and handles encoding detection.

Stage 4: AST Traversal and Normalisation

The ASTTraverser class in graphify/src/ast/ast_traverser.py executes a depth-first walk of the Tree-sitter tree. It creates a Graphify Node for every Tree-sitter named node (functions, classes, statements), generating deterministic Node IDs by hashing the node's start byte, end byte, and type. The traverser also creates edges representing parent-child relationships from the Tree-sitter hierarchy.

Stage 5: Edge Enrichment

After extracting the raw tree structure, graphify/src/ast/edge_enricher.py augments the graph with semantic edge stubs. These placeholders for calls, references, and inherits relationships are later populated by language-specific analyzers during the analysis phase.

Stage 6: Persistence

The constructed nodes and edges are written to the backing store through graphify/src/storage/graph_store.py. Graphify supports both Neo4j and SQLite backends, batch-inserting the AST fragments to enable fast querying for downstream analyses.

Stage 7: Caching

To avoid re-parsing unchanged files, the pipeline checks graphify/src/cache/ast_cache.py for existing ASTs. Parsed trees are serialized to disk in /tmp/graphify/cache/, keyed by content hash, allowing subsequent runs to skip redundant parsing work.

Integration Architecture and Entry Points

Understanding how these stages connect reveals the pipeline's orchestration layer.

Entry Point and File Routing

The process begins in graphify/main.py, which iterates over the repository files and dispatches each to the appropriate handler. The FileRouter class in graphify/src/router/file_router.py determines whether a file requires Tree-sitter parsing (for code) or LLM-based semantic extraction (for documentation and images).

Core Parser Classes

The TreeSitterParser class (in tree_sitter_parser.py) creates the initial Tree object, then hands it to the ASTTraverser (in ast_traverser.py). The traverser emits Graphify-specific objects: ASTNode for vertices and ASTEdge for relationships, creating a clean abstraction layer over the raw Tree-sitter output.

Deterministic Identity Generation

Node IDs are generated in graphify/src/utils/hash_utils.py using deterministic hashing of node attributes. This guarantees that the same source file always yields the same graph fragment, allowing Graphify to merge multiple extraction runs without duplication or identity conflicts.

Why Tree-sitter Powers the Pipeline

Graphify selected Tree-sitter as its parsing foundation for three critical reasons:

  • Speed: Tree-sitter parses locally in less than 10 milliseconds for most source files, keeping extraction costs near zero.
  • Language Coverage: The uv.lock file defines 29 supported languages, from JavaScript to Rust, each loaded as a separate compiled grammar.
  • Deterministic Structure: Unlike regular expression-based tokenizers, Tree-sitter guarantees a full, correct syntactic tree that Graphify can translate directly into graph edges without heuristic parsing.

Practical Usage Examples

Below are concrete implementations showing how to interact with the pipeline programmatically.

This Python example demonstrates the high-level router interface:


# example.py – parsing a JavaScript file and inserting its AST into the graph store

from graphify.src.router.file_router import FileRouter
from graphify.src.storage.graph_store import GraphStore

router = FileRouter()
store = GraphStore()

js_path = "src/app/main.js"
if router.is_code(js_path):
    ast = router.parse_with_tree_sitter(js_path)   # returns Graphify AST object

    store.save_ast(ast)                            # persists nodes & edges

For lower-level control, you can interact directly with the parser and traverser:

// ts_example.ts – manual low‑level usage of the Tree‑sitter parser
import { TreeSitterParser } from "graphify/src/parsers/tree_sitter_parser";
import { ASTTraverser } from "graphify/src/ast/ast_traverser";

const source = await Deno.readTextFile("src/utils/helpers.ts");
const parser = new TreeSitterParser("typescript");
const tree = parser.parse(source);
const traverser = new ASTTraverser();
const graphAst = traverser.traverse(tree);
console.log(graphAst.nodes.length, "nodes extracted");

Summary

  • Graphify's pipeline consists of seven stages: language detection, parser loading, source-file parsing, AST traversal, edge enrichment, persistence, and caching.
  • The TreeSitterParser and ASTTraverser classes in graphify/src/parsers/ and graphify/src/ast/ form the core extraction engine.
  • Deterministic hashing in graphify/src/utils/hash_utils.py ensures identical files produce identical graph fragments.
  • Parsed ASTs are cached in /tmp/graphify/cache/ via graphify/src/cache/ast_cache.py to eliminate redundant processing.
  • The system supports 29 languages defined in uv.lock and persists results to Neo4j or SQLite via graphify/src/storage/graph_store.py.

Frequently Asked Questions

How does Graphify ensure deterministic AST extraction across multiple runs?

Graphify generates node IDs using deterministic hashes of each node's start byte, end byte, and type in graphify/src/utils/hash_utils.py. This guarantees that identical source files always produce identical graph fragments, enabling safe merging of multiple extraction runs without duplication.

What happens when Graphify encounters a file type it doesn't support?

The FileRouter class in graphify/src/router/file_router.py routes unsupported code files and non-code assets (like documentation or images) to an LLM-based semantic extractor rather than the Tree-sitter pipeline. This ensures comprehensive coverage even for languages not in the 29 supported Tree-sitter grammars listed in uv.lock.

Where does Graphify store cached ASTs to avoid re-parsing?

The system serializes parsed ASTs to disk in /tmp/graphify/cache/ using the ASTCache module in graphify/src/cache/ast_cache.py. Each cache entry is keyed by content hash, allowing subsequent runs to skip parsing for any file that hasn't changed since the last extraction.

Which storage backends does Graphify support for the extracted graph data?

According to the source code in graphify/src/storage/graph_store.py, Graphify persists nodes and edges to either Neo4j (for production graph databases) or SQLite (for local development). This dual-backend approach provides flexibility for different deployment scenarios.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →