How to Add Support for a New Programming Language in Graphify
To add support for a new programming language in Graphify, you must create a Tree-Sitter-based extractor module in graphify/extractors/<lang>.py, register it in graphify/extractors/__init__.py and models.py, and declare the tree-sitter-<lang> dependency in pyproject.toml to enable parsing and graph generation.
Graphify is an open-source code intelligence tool that converts source code into knowledge graphs by extracting symbols and relationships using language-specific parsers. If you need to add support for a new programming language in Graphify, you will implement a Tree-Sitter extractor that transforms Abstract Syntax Tree (AST) nodes into Graphify's standardized node and edge format. This guide references the actual source files in the Graphify-Labs/graphify repository to show you exactly how to integrate a new language.
Step 1: Create the Language-Specific Extractor
Create a new Python module at graphify/extractors/<lang>.py that implements the extraction logic. This module must parse source files using a Tree-Sitter grammar and emit nodes and edges following Graphify's JSON schema.
The Extractor Template
Here is the complete template you should adapt for your target language:
"""<Lang> extractor – parses <lang> files and emits Graphify nodes/edges."""
from __future__ import annotations
from pathlib import Path
from graphify.extractors.base import (
_LANGUAGE_BUILTIN_GLOBALS,
_file_stem,
_make_id,
_read_text,
)
# ----------------------------------------------------------------------
# Helper: collect type references (adjust for your language’s AST)
def _<lang>_collect_type_refs(node, source: bytes, generic: bool,
out: list[tuple[str, str]]) -> None:
"""Walk a <Lang> type expression and append (name, role) tuples."""
if node is None:
return
# Example – handle simple identifier & generic arguments
if node.type == "type_identifier":
text = _read_text(node, source)
if text:
out.append((text, "generic_arg" if generic else "type"))
elif node.type == "generic_type":
# Recurse into generic arguments
for child in node.children:
if child.is_named:
_<lang>_collect_type_refs(child, source, True, out)
else:
for child in node.children:
if child.is_named:
_<lang>_collect_type_refs(child, source, generic, out)
# ----------------------------------------------------------------------
def extract_<lang>(path: Path) -> dict:
"""Extract top‑level symbols from a *. <lang> file."""
try:
import tree_sitter_<lang> as ts_lang # ← Tree‑Sitter binding
from tree_sitter import Language, Parser
except ImportError:
# Graphify will report the missing parser in its UI
return {"nodes": [], "edges": [], "error": "tree‑sitter-<lang> not installed"}
# Parse the file
language = Language(ts_lang.language())
parser = Parser(language)
source = path.read_bytes()
tree = parser.parse(source)
root = tree.root_node
stem = _file_stem(path)
str_path = str(path)
nodes, edges, seen = [], [], set()
# ------------------------------------------------------------------
def add_node(nid: str, label: str, line: int) -> None:
if nid not in seen:
seen.add(nid)
nodes.append({
"id": nid,
"label": label,
"file_type": "code",
"source_file": str_path,
"source_location": f"L{line}",
})
def add_edge(src: str, tgt: str, rel: str, line: int,
context: str | None = None) -> None:
edge = {
"source": src,
"target": tgt,
"relation": rel,
"confidence": "EXTRACTED",
"source_file": str_path,
"source_location": f"L{line}",
"weight": 1.0,
}
if context:
edge["context"] = context
edges.append(edge)
# ------------------------------------------------------------------
file_nid = _make_id(str(path))
add_node(file_nid, path.name, 1)
# Walk the AST and emit symbols – adapt to your language’s node types
def walk(node, parent_impl: str | None = None) -> None:
if node.type == "function_item":
name_node = node.child_by_field_name("name")
if name_node:
func_name = _read_text(name_node, source)
line = node.start_point[0] + 1
func_nid = _make_id(stem, func_name)
add_node(func_nid, f"{func_name}()", line)
add_edge(file_nid, func_nid, "contains", line)
# Export parameter & return‑type references
# (reuse the helper above)
# …
# Add similar branches for structs, enums, etc.
for child in node.children:
if child.is_named:
walk(child, parent_impl)
walk(root)
return {"nodes": nodes, "edges": edges}
Key Implementation Details
- Import the Tree-Sitter binding inside a
try/exceptblock—Graphify gracefully degrades when the parser is missing by returning an error payload. - Use base utilities from
graphify.extractors.baseto keep IDs stable across runs:_file_stemextracts the filename without extension,_make_idgenerates unique identifiers, and_read_textsafely extracts node text. - Emit nodes and edges via the
add_nodeandadd_edgehelpers—Graphify expects the same JSON schema that existing extractors use. - Reference implementation: See
graphify/extractors/rust.pyin the repository for a fully-featured example.
Step 2: Register the Extractor in Graphify
You must wire the new extractor into three specific files to make it discoverable by the framework.
Update the Extractors Module
Edit graphify/extractors/__init__.py to import the new extractor function:
from graphify.extractors.<lang> import extract_<lang>
This import tells Graphify that the extractor exists and is available for the file extensions you will configure next.
Add Language Configuration
Edit graphify/extractors/models.py to add a LanguageConfig entry that maps file extensions to your extractor:
LanguageConfig(name="<Lang>", ts_module="tree_sitter_<lang>", extensions=[".<ext>"])
The ts_module parameter must match the Python package name of the Tree-Sitter binding (e.g., tree_sitter_rust for Rust).
Declare Dependencies
Edit pyproject.toml to declare the Tree-Sitter grammar as an optional dependency. This allows users to install support for your language specifically:
[project.optional-dependencies]
<lang> = ["tree-sitter-<lang>==<version>"]
Users can then install your language support with pip install graphify[<lang>].
Step 3: Provide the Tree-Sitter Grammar
Graphify relies on Tree-Sitter parsers to generate ASTs. You must ensure the tree-sitter-<lang> package is available in the environment.
- Install the grammar via pip:
pip install tree-sitter-<lang> - Update CI configuration (optional but recommended): Add the new parser to
.github/workflows/ci.ymlso the test suite validates your extractor. - Handle missing parsers: When the parser is unavailable, your extractor returns
{"error": "tree‑sitter-<lang> not installed"}, and Graphify's UI will surface the missing dependency without crashing.
Step 4: Test Your New Extractor
Create a minimal test in tests/test_<lang>_extraction.py to verify that your extractor correctly parses source files and generates valid graph nodes:
def test_<lang>_extraction(tmp_path):
src = tmp_path / "example.<ext>"
src.write_text("fn hello() { }") # Adjust to your language syntax
result = graphify.extract.extract_<lang>(src)
assert any(node["label"] == "hello()" for node in result["nodes"])
Running pytest validates that the extractor conforms to Graphify's schema and correctly identifies symbols from the target language.
Summary
- Create
graphify/extractors/<lang>.pyimplementingextract_<lang>()with Tree-Sitter AST walking and Graphify node/edge emission. - Register the extractor in
graphify/extractors/__init__.pyby importing the extraction function. - Configure the language in
graphify/extractors/models.pyusingLanguageConfigwith the correctts_moduleand file extensions. - Declare the
tree-sitter-<lang>dependency inpyproject.tomlas an optional extra. - Test the implementation in
tests/test_<lang>_extraction.pyto verify symbol extraction and schema compliance.
Frequently Asked Questions
What is the Tree-Sitter requirement for adding a new language?
Graphify uses Tree-Sitter parsers to generate Abstract Syntax Trees (ASTs) from source code. You must provide a Python binding package named tree-sitter-<lang> that exposes a language() function compatible with the tree_sitter library. Without this binding, Graphify cannot parse the source code into an AST for your extractor to traverse.
How do I handle missing Tree-Sitter parsers gracefully?
Wrap the Tree-Sitter import in a try/except block inside your extract_<lang>() function. If the import fails, return a dictionary with an "error" key containing "tree‑sitter-<lang> not installed". Graphify's engine detects this error payload and displays a missing dependency message in the UI without halting the entire extraction process.
Where can I find a reference implementation for a Graphify extractor?
The graphify/extractors/rust.py file in the Graphify-Labs/graphify repository serves as the canonical reference implementation. It demonstrates advanced techniques including type reference collection, handling of generic types, and proper usage of base utilities like _make_id and _read_text for generating stable graph identifiers.
What file extensions do I need to modify to register a new language?
You must modify exactly three files: graphify/extractors/__init__.py to import the extractor function, graphify/extractors/models.py to define the LanguageConfig mapping file extensions to the extractor, and pyproject.toml to declare the optional dependency. Additionally, you should create the extractor file itself at graphify/extractors/<lang>.py.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →