# How to Add Support for a New Programming Language in Graphify

> Learn to add new programming language support to Graphify by creating a Tree-Sitter extractor module. Follow our guide to integrate your language seamlessly.

- Repository: [Graphify Labs/graphify](https://github.com/Graphify-Labs/graphify)
- Tags: how-to-guide
- Published: 2026-07-16

---

**To add support for a new programming language in Graphify, you must create a Tree-Sitter-based extractor module in `graphify/extractors/<lang>.py`, register it in [`graphify/extractors/__init__.py`](https://github.com/Graphify-Labs/graphify/blob/main/graphify/extractors/__init__.py) and [`models.py`](https://github.com/Graphify-Labs/graphify/blob/main/models.py), and declare the `tree-sitter-<lang>` dependency in [`pyproject.toml`](https://github.com/Graphify-Labs/graphify/blob/main/pyproject.toml) to enable parsing and graph generation.**

Graphify is an open-source code intelligence tool that converts source code into knowledge graphs by extracting symbols and relationships using language-specific parsers. If you need to add support for a new programming language in Graphify, you will implement a **Tree-Sitter extractor** that transforms Abstract Syntax Tree (AST) nodes into Graphify's standardized node and edge format. This guide references the actual source files in the `Graphify-Labs/graphify` repository to show you exactly how to integrate a new language.

## Step 1: Create the Language-Specific Extractor

Create a new Python module at `graphify/extractors/<lang>.py` that implements the extraction logic. This module must parse source files using a Tree-Sitter grammar and emit nodes and edges following Graphify's JSON schema.

### The Extractor Template

Here is the complete template you should adapt for your target language:

```python
"""<Lang> extractor – parses <lang> files and emits Graphify nodes/edges."""
from __future__ import annotations

from pathlib import Path
from graphify.extractors.base import (
    _LANGUAGE_BUILTIN_GLOBALS,
    _file_stem,
    _make_id,
    _read_text,
)

# ----------------------------------------------------------------------

# Helper: collect type references (adjust for your language’s AST)

def _<lang>_collect_type_refs(node, source: bytes, generic: bool,
                             out: list[tuple[str, str]]) -> None:
    """Walk a <Lang> type expression and append (name, role) tuples."""
    if node is None:
        return
    # Example – handle simple identifier & generic arguments

    if node.type == "type_identifier":
        text = _read_text(node, source)
        if text:
            out.append((text, "generic_arg" if generic else "type"))
    elif node.type == "generic_type":
        # Recurse into generic arguments

        for child in node.children:
            if child.is_named:
                _<lang>_collect_type_refs(child, source, True, out)
    else:
        for child in node.children:
            if child.is_named:
                _<lang>_collect_type_refs(child, source, generic, out)

# ----------------------------------------------------------------------

def extract_<lang>(path: Path) -> dict:
    """Extract top‑level symbols from a *. <lang> file."""
    try:
        import tree_sitter_<lang> as ts_lang               # ← Tree‑Sitter binding

        from tree_sitter import Language, Parser
    except ImportError:
        # Graphify will report the missing parser in its UI

        return {"nodes": [], "edges": [], "error": "tree‑sitter-<lang> not installed"}

    # Parse the file

    language = Language(ts_lang.language())
    parser = Parser(language)
    source = path.read_bytes()
    tree = parser.parse(source)
    root = tree.root_node

    stem = _file_stem(path)
    str_path = str(path)
    nodes, edges, seen = [], [], set()

    # ------------------------------------------------------------------

    def add_node(nid: str, label: str, line: int) -> None:
        if nid not in seen:
            seen.add(nid)
            nodes.append({
                "id": nid,
                "label": label,
                "file_type": "code",
                "source_file": str_path,
                "source_location": f"L{line}",
            })

    def add_edge(src: str, tgt: str, rel: str, line: int,
                 context: str | None = None) -> None:
        edge = {
            "source": src,
            "target": tgt,
            "relation": rel,
            "confidence": "EXTRACTED",
            "source_file": str_path,
            "source_location": f"L{line}",
            "weight": 1.0,
        }
        if context:
            edge["context"] = context
        edges.append(edge)

    # ------------------------------------------------------------------

    file_nid = _make_id(str(path))
    add_node(file_nid, path.name, 1)

    # Walk the AST and emit symbols – adapt to your language’s node types

    def walk(node, parent_impl: str | None = None) -> None:
        if node.type == "function_item":
            name_node = node.child_by_field_name("name")
            if name_node:
                func_name = _read_text(name_node, source)
                line = node.start_point[0] + 1
                func_nid = _make_id(stem, func_name)
                add_node(func_nid, f"{func_name}()", line)
                add_edge(file_nid, func_nid, "contains", line)
                # Export parameter & return‑type references

                # (reuse the helper above)

                # …

        # Add similar branches for structs, enums, etc.

        for child in node.children:
            if child.is_named:
                walk(child, parent_impl)

    walk(root)

    return {"nodes": nodes, "edges": edges}

```

### Key Implementation Details

- **Import the Tree-Sitter binding** inside a `try/except` block—Graphify gracefully degrades when the parser is missing by returning an error payload.
- **Use base utilities** from `graphify.extractors.base` to keep IDs stable across runs: `_file_stem` extracts the filename without extension, `_make_id` generates unique identifiers, and `_read_text` safely extracts node text.
- **Emit nodes and edges** via the `add_node` and `add_edge` helpers—Graphify expects the same JSON schema that existing extractors use.
- **Reference implementation**: See [`graphify/extractors/rust.py`](https://github.com/Graphify-Labs/graphify/blob/main/graphify/extractors/rust.py) in the repository for a fully-featured example.

## Step 2: Register the Extractor in Graphify

You must wire the new extractor into three specific files to make it discoverable by the framework.

### Update the Extractors Module

Edit [`graphify/extractors/__init__.py`](https://github.com/Graphify-Labs/graphify/blob/main/graphify/extractors/__init__.py) to import the new extractor function:

```python
from graphify.extractors.<lang> import extract_<lang>

```

This import tells Graphify that the extractor exists and is available for the file extensions you will configure next.

### Add Language Configuration

Edit [`graphify/extractors/models.py`](https://github.com/Graphify-Labs/graphify/blob/main/graphify/extractors/models.py) to add a **LanguageConfig** entry that maps file extensions to your extractor:

```python
LanguageConfig(name="<Lang>", ts_module="tree_sitter_<lang>", extensions=[".<ext>"])

```

The `ts_module` parameter must match the Python package name of the Tree-Sitter binding (e.g., `tree_sitter_rust` for Rust).

### Declare Dependencies

Edit [`pyproject.toml`](https://github.com/Graphify-Labs/graphify/blob/main/pyproject.toml) to declare the Tree-Sitter grammar as an optional dependency. This allows users to install support for your language specifically:

```toml
[project.optional-dependencies]
<lang> = ["tree-sitter-<lang>==<version>"]

```

Users can then install your language support with `pip install graphify[<lang>]`.

## Step 3: Provide the Tree-Sitter Grammar

Graphify relies on Tree-Sitter parsers to generate ASTs. You must ensure the `tree-sitter-<lang>` package is available in the environment.

1. **Install the grammar** via pip: `pip install tree-sitter-<lang>`
2. **Update CI configuration** (optional but recommended): Add the new parser to [`.github/workflows/ci.yml`](https://github.com/Graphify-Labs/graphify/blob/main/.github/workflows/ci.yml) so the test suite validates your extractor.
3. **Handle missing parsers**: When the parser is unavailable, your extractor returns `{"error": "tree‑sitter-<lang> not installed"}`, and Graphify's UI will surface the missing dependency without crashing.

## Step 4: Test Your New Extractor

Create a minimal test in `tests/test_<lang>_extraction.py` to verify that your extractor correctly parses source files and generates valid graph nodes:

```python
def test_<lang>_extraction(tmp_path):
    src = tmp_path / "example.<ext>"
    src.write_text("fn hello() { }")          # Adjust to your language syntax

    result = graphify.extract.extract_<lang>(src)
    assert any(node["label"] == "hello()" for node in result["nodes"])

```

Running `pytest` validates that the extractor conforms to Graphify's schema and correctly identifies symbols from the target language.

## Summary

- **Create** `graphify/extractors/<lang>.py` implementing `extract_<lang>()` with Tree-Sitter AST walking and Graphify node/edge emission.
- **Register** the extractor in [`graphify/extractors/__init__.py`](https://github.com/Graphify-Labs/graphify/blob/main/graphify/extractors/__init__.py) by importing the extraction function.
- **Configure** the language in [`graphify/extractors/models.py`](https://github.com/Graphify-Labs/graphify/blob/main/graphify/extractors/models.py) using `LanguageConfig` with the correct `ts_module` and file extensions.
- **Declare** the `tree-sitter-<lang>` dependency in [`pyproject.toml`](https://github.com/Graphify-Labs/graphify/blob/main/pyproject.toml) as an optional extra.
- **Test** the implementation in `tests/test_<lang>_extraction.py` to verify symbol extraction and schema compliance.

## Frequently Asked Questions

### What is the Tree-Sitter requirement for adding a new language?

Graphify uses Tree-Sitter parsers to generate Abstract Syntax Trees (ASTs) from source code. You must provide a Python binding package named `tree-sitter-<lang>` that exposes a `language()` function compatible with the `tree_sitter` library. Without this binding, Graphify cannot parse the source code into an AST for your extractor to traverse.

### How do I handle missing Tree-Sitter parsers gracefully?

Wrap the Tree-Sitter import in a `try/except` block inside your `extract_<lang>()` function. If the import fails, return a dictionary with an `"error"` key containing `"tree‑sitter-<lang> not installed"`. Graphify's engine detects this error payload and displays a missing dependency message in the UI without halting the entire extraction process.

### Where can I find a reference implementation for a Graphify extractor?

The [`graphify/extractors/rust.py`](https://github.com/Graphify-Labs/graphify/blob/main/graphify/extractors/rust.py) file in the `Graphify-Labs/graphify` repository serves as the canonical reference implementation. It demonstrates advanced techniques including type reference collection, handling of generic types, and proper usage of base utilities like `_make_id` and `_read_text` for generating stable graph identifiers.

### What file extensions do I need to modify to register a new language?

You must modify exactly three files: [`graphify/extractors/__init__.py`](https://github.com/Graphify-Labs/graphify/blob/main/graphify/extractors/__init__.py) to import the extractor function, [`graphify/extractors/models.py`](https://github.com/Graphify-Labs/graphify/blob/main/graphify/extractors/models.py) to define the `LanguageConfig` mapping file extensions to the extractor, and [`pyproject.toml`](https://github.com/Graphify-Labs/graphify/blob/main/pyproject.toml) to declare the optional dependency. Additionally, you should create the extractor file itself at `graphify/extractors/<lang>.py`.