How to Add Support for a New Programming Language in Graphify

To add support for a new programming language in Graphify, you must create a Tree-Sitter-based extractor module in graphify/extractors/<lang>.py, register it in graphify/extractors/__init__.py and models.py, and declare the tree-sitter-<lang> dependency in pyproject.toml to enable parsing and graph generation.

Graphify is an open-source code intelligence tool that converts source code into knowledge graphs by extracting symbols and relationships using language-specific parsers. If you need to add support for a new programming language in Graphify, you will implement a Tree-Sitter extractor that transforms Abstract Syntax Tree (AST) nodes into Graphify's standardized node and edge format. This guide references the actual source files in the Graphify-Labs/graphify repository to show you exactly how to integrate a new language.

Step 1: Create the Language-Specific Extractor

Create a new Python module at graphify/extractors/<lang>.py that implements the extraction logic. This module must parse source files using a Tree-Sitter grammar and emit nodes and edges following Graphify's JSON schema.

The Extractor Template

Here is the complete template you should adapt for your target language:

"""<Lang> extractor – parses <lang> files and emits Graphify nodes/edges."""
from __future__ import annotations

from pathlib import Path
from graphify.extractors.base import (
    _LANGUAGE_BUILTIN_GLOBALS,
    _file_stem,
    _make_id,
    _read_text,
)

# ----------------------------------------------------------------------

# Helper: collect type references (adjust for your language’s AST)

def _<lang>_collect_type_refs(node, source: bytes, generic: bool,
                             out: list[tuple[str, str]]) -> None:
    """Walk a <Lang> type expression and append (name, role) tuples."""
    if node is None:
        return
    # Example – handle simple identifier & generic arguments

    if node.type == "type_identifier":
        text = _read_text(node, source)
        if text:
            out.append((text, "generic_arg" if generic else "type"))
    elif node.type == "generic_type":
        # Recurse into generic arguments

        for child in node.children:
            if child.is_named:
                _<lang>_collect_type_refs(child, source, True, out)
    else:
        for child in node.children:
            if child.is_named:
                _<lang>_collect_type_refs(child, source, generic, out)

# ----------------------------------------------------------------------

def extract_<lang>(path: Path) -> dict:
    """Extract top‑level symbols from a *. <lang> file."""
    try:
        import tree_sitter_<lang> as ts_lang               # ← Tree‑Sitter binding

        from tree_sitter import Language, Parser
    except ImportError:
        # Graphify will report the missing parser in its UI

        return {"nodes": [], "edges": [], "error": "tree‑sitter-<lang> not installed"}

    # Parse the file

    language = Language(ts_lang.language())
    parser = Parser(language)
    source = path.read_bytes()
    tree = parser.parse(source)
    root = tree.root_node

    stem = _file_stem(path)
    str_path = str(path)
    nodes, edges, seen = [], [], set()

    # ------------------------------------------------------------------

    def add_node(nid: str, label: str, line: int) -> None:
        if nid not in seen:
            seen.add(nid)
            nodes.append({
                "id": nid,
                "label": label,
                "file_type": "code",
                "source_file": str_path,
                "source_location": f"L{line}",
            })

    def add_edge(src: str, tgt: str, rel: str, line: int,
                 context: str | None = None) -> None:
        edge = {
            "source": src,
            "target": tgt,
            "relation": rel,
            "confidence": "EXTRACTED",
            "source_file": str_path,
            "source_location": f"L{line}",
            "weight": 1.0,
        }
        if context:
            edge["context"] = context
        edges.append(edge)

    # ------------------------------------------------------------------

    file_nid = _make_id(str(path))
    add_node(file_nid, path.name, 1)

    # Walk the AST and emit symbols – adapt to your language’s node types

    def walk(node, parent_impl: str | None = None) -> None:
        if node.type == "function_item":
            name_node = node.child_by_field_name("name")
            if name_node:
                func_name = _read_text(name_node, source)
                line = node.start_point[0] + 1
                func_nid = _make_id(stem, func_name)
                add_node(func_nid, f"{func_name}()", line)
                add_edge(file_nid, func_nid, "contains", line)
                # Export parameter & return‑type references

                # (reuse the helper above)

                # …

        # Add similar branches for structs, enums, etc.

        for child in node.children:
            if child.is_named:
                walk(child, parent_impl)

    walk(root)

    return {"nodes": nodes, "edges": edges}

Key Implementation Details

  • Import the Tree-Sitter binding inside a try/except block—Graphify gracefully degrades when the parser is missing by returning an error payload.
  • Use base utilities from graphify.extractors.base to keep IDs stable across runs: _file_stem extracts the filename without extension, _make_id generates unique identifiers, and _read_text safely extracts node text.
  • Emit nodes and edges via the add_node and add_edge helpers—Graphify expects the same JSON schema that existing extractors use.
  • Reference implementation: See graphify/extractors/rust.py in the repository for a fully-featured example.

Step 2: Register the Extractor in Graphify

You must wire the new extractor into three specific files to make it discoverable by the framework.

Update the Extractors Module

Edit graphify/extractors/__init__.py to import the new extractor function:

from graphify.extractors.<lang> import extract_<lang>

This import tells Graphify that the extractor exists and is available for the file extensions you will configure next.

Add Language Configuration

Edit graphify/extractors/models.py to add a LanguageConfig entry that maps file extensions to your extractor:

LanguageConfig(name="<Lang>", ts_module="tree_sitter_<lang>", extensions=[".<ext>"])

The ts_module parameter must match the Python package name of the Tree-Sitter binding (e.g., tree_sitter_rust for Rust).

Declare Dependencies

Edit pyproject.toml to declare the Tree-Sitter grammar as an optional dependency. This allows users to install support for your language specifically:

[project.optional-dependencies]
<lang> = ["tree-sitter-<lang>==<version>"]

Users can then install your language support with pip install graphify[<lang>].

Step 3: Provide the Tree-Sitter Grammar

Graphify relies on Tree-Sitter parsers to generate ASTs. You must ensure the tree-sitter-<lang> package is available in the environment.

  1. Install the grammar via pip: pip install tree-sitter-<lang>
  2. Update CI configuration (optional but recommended): Add the new parser to .github/workflows/ci.yml so the test suite validates your extractor.
  3. Handle missing parsers: When the parser is unavailable, your extractor returns {"error": "tree‑sitter-<lang> not installed"}, and Graphify's UI will surface the missing dependency without crashing.

Step 4: Test Your New Extractor

Create a minimal test in tests/test_<lang>_extraction.py to verify that your extractor correctly parses source files and generates valid graph nodes:

def test_<lang>_extraction(tmp_path):
    src = tmp_path / "example.<ext>"
    src.write_text("fn hello() { }")          # Adjust to your language syntax

    result = graphify.extract.extract_<lang>(src)
    assert any(node["label"] == "hello()" for node in result["nodes"])

Running pytest validates that the extractor conforms to Graphify's schema and correctly identifies symbols from the target language.

Summary

  • Create graphify/extractors/<lang>.py implementing extract_<lang>() with Tree-Sitter AST walking and Graphify node/edge emission.
  • Register the extractor in graphify/extractors/__init__.py by importing the extraction function.
  • Configure the language in graphify/extractors/models.py using LanguageConfig with the correct ts_module and file extensions.
  • Declare the tree-sitter-<lang> dependency in pyproject.toml as an optional extra.
  • Test the implementation in tests/test_<lang>_extraction.py to verify symbol extraction and schema compliance.

Frequently Asked Questions

What is the Tree-Sitter requirement for adding a new language?

Graphify uses Tree-Sitter parsers to generate Abstract Syntax Trees (ASTs) from source code. You must provide a Python binding package named tree-sitter-<lang> that exposes a language() function compatible with the tree_sitter library. Without this binding, Graphify cannot parse the source code into an AST for your extractor to traverse.

How do I handle missing Tree-Sitter parsers gracefully?

Wrap the Tree-Sitter import in a try/except block inside your extract_<lang>() function. If the import fails, return a dictionary with an "error" key containing "tree‑sitter-<lang> not installed". Graphify's engine detects this error payload and displays a missing dependency message in the UI without halting the entire extraction process.

Where can I find a reference implementation for a Graphify extractor?

The graphify/extractors/rust.py file in the Graphify-Labs/graphify repository serves as the canonical reference implementation. It demonstrates advanced techniques including type reference collection, handling of generic types, and proper usage of base utilities like _make_id and _read_text for generating stable graph identifiers.

What file extensions do I need to modify to register a new language?

You must modify exactly three files: graphify/extractors/__init__.py to import the extractor function, graphify/extractors/models.py to define the LanguageConfig mapping file extensions to the extractor, and pyproject.toml to declare the optional dependency. Additionally, you should create the extractor file itself at graphify/extractors/<lang>.py.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →