# Development Roadmap for Graphify: Pipeline Stability, Language Expansion, and Analytics

> Explore the Graphify development roadmap. Discover improvements in pipeline stability, new language support with tree-sitter, and advanced graph analytics for scalable insights.

- Repository: [Graphify Labs/graphify](https://github.com/Graphify-Labs/graphify)
- Tags: development-roadmap
- Published: 2026-07-19

---

**Graphify's development roadmap prioritizes stabilizing the core extraction pipeline through critical bug fixes, expanding language coverage via modular tree-sitter extractors, and scaling graph analytics with parallel clustering and enhanced export formats.**

The development roadmap for Graphify (Graphify-Labs/graphify) focuses on hardening the code-to-graph extraction engine while broadening support for modern programming languages. According to the project's [`ARCHITECTURE.md`](https://github.com/Graphify-Labs/graphify/blob/main/ARCHITECTURE.md) and test suites, the maintainers are actively addressing five documented bugs in the extraction stages and standardizing the process for adding new language parsers. This roadmap reflects an open-ended, community-driven approach where contributors can prioritize new extractors or performance improvements based on demand.

## Stabilizing the Core Extraction Pipeline

The immediate priority in the Graphify roadmap is **production stability** for the extraction pipeline. The architecture follows a strict four-stage flow defined in [`ARCHITECTURE.md`](https://github.com/Graphify-Labs/graphify/blob/main/ARCHITECTURE.md): **detect** → **extract** → **build** → **analyze**.

Critical to this phase is the "roadmap bug-fixes" test case found in [`tests/test_dart.py`](https://github.com/Graphify-Labs/graphify/blob/main/tests/test_dart.py), which explicitly tracks five known bugs (labeled A through E) affecting the extraction stages. These fixes target edge cases in AST parsing and edge inference that currently impact the reliability of the `extract()` dispatcher in [`graphify/extract.py`](https://github.com/Graphify-Labs/graphify/blob/main/graphify/extract.py). Resolving these issues ensures that the **detect** phase (handled by [`graphify/detect.py`](https://github.com/Graphify-Labs/graphify/blob/main/graphify/detect.py)) and the **extract** phase produce consistent node and edge sets before the graph construction layer processes them.

## Expanding Language Coverage and Extractor Quality

The second pillar of the roadmap emphasizes **language ecosystem growth**. Graphify currently ships with **36 tree-sitter grammars** and provides a documented, repeatable process in [`ARCHITECTURE.md`](https://github.com/Graphify-Labs/graphify/blob/main/ARCHITECTURE.md) for adding support for additional languages.

To implement a new language extractor, contributors must:

- Create an `extract_<lang>` function in [`graphify/extract.py`](https://github.com/Graphify-Labs/graphify/blob/main/graphify/extract.py) that returns a `{"nodes": [], "edges": []}` dictionary
- Register the file extension in [`graphify/detect.py`](https://github.com/Graphify-Labs/graphify/blob/main/graphify/detect.py) (updating `CODE_EXTENSIONS`) and [`graphify/watch.py`](https://github.com/Graphify-Labs/graphify/blob/main/graphify/watch.py) (updating `_WATCHED_EXTENSIONS`)
- Add the tree-sitter grammar as an optional dependency in [`pyproject.toml`](https://github.com/Graphify-Labs/graphify/blob/main/pyproject.toml)
- Provide test fixtures in `tests/fixtures/` and unit tests in [`tests/test_languages.py`](https://github.com/Graphify-Labs/graphify/blob/main/tests/test_languages.py)

Target languages explicitly mentioned for future inclusion include **Rust** and **Kotlin**, though the modular architecture supports any language with a tree-sitter grammar. The optional extras system allows users to install language-specific parsers on demand without bloating the core installation.

## Improving Graph Analytics and Scalability

The third strategic focus addresses **performance and export capabilities** for large-scale codebases. Currently, community detection and "god-node" identification rely on Leiden clustering algorithms (referenced in [`benchmark.py`](https://github.com/Graphify-Labs/graphify/blob/main/benchmark.py)) and LLM-free heuristics with three confidence levels: **EXTRACTED**, **INFERRED**, and **AMBIGUOUS**.

Future milestones include:

- **Parallel processing**: The `graphify cluster-only` command already accepts `--max-concurrency` and `--resolution` flags, indicating active development toward faster, multi-threaded community detection
- **Intelligent deduplication**: Automatic merging of duplicate nodes with improved confidence tagging
- **Database integration**: Native export formats for **Neo4j**, **FalkorDB**, and **SVG** visualization
- **MCP server support**: Real-time querying integration with Model Context Protocol servers for AI-assisted code analysis

## How to Contribute: Adding a New Language Extractor

The following example demonstrates the roadmap's standardized process for adding a "FooLang" extractor, following the implementation pattern established in [`graphify/extract.py`](https://github.com/Graphify-Labs/graphify/blob/main/graphify/extract.py):

```python

# graphify/extract.py

from pathlib import Path
import tree_sitter  # ← the parser for the new language

def extract_foolang(path: Path) -> dict:
    """Parse a FooLang file and return a nodes/edges dict."""
    source = path.read_text()
    parser = tree_sitter.Parser()
    parser.set_language(tree_sitter.Language('build/my-languages.so', 'foolang'))
    tree = parser.parse(bytes(source, "utf8"))

    # Walk the AST, collect nodes and explicit edges.

    nodes, edges = [], []
    # ... (AST walking logic omitted for brevity) ...

    # Second-pass call-graph to infer additional `calls` edges.

    # (See existing extractors for reference.)

    return {"nodes": nodes, "edges": edges}

# Dispatch logic (already present in the file):

def extract(path: Path) -> dict:
    if path.suffix == ".foo":
        return extract_foolang(path)
    # existing language dispatches …

```

To complete the integration, update `CODE_EXTENSIONS` in [`graphify/detect.py`](https://github.com/Graphify-Labs/graphify/blob/main/graphify/detect.py) and `_WATCHED_EXTENSIONS` in [`graphify/watch.py`](https://github.com/Graphify-Labs/graphify/blob/main/graphify/watch.py) to include `.foo`, then add the grammar dependency to [`pyproject.toml`](https://github.com/Graphify-Labs/graphify/blob/main/pyproject.toml).

## Summary

- **Immediate priorities**: Resolving bugs A-E in the extraction pipeline, documented in [`tests/test_dart.py`](https://github.com/Graphify-Labs/graphify/blob/main/tests/test_dart.py), to stabilize the detect-extract-build-analyze flow
- **Language expansion**: Supporting additional languages like Rust and Kotlin through the modular `extract_<lang>` pattern and optional dependency management
- **Scalability improvements**: Implementing parallel community detection (`--max-concurrency`), expanding export formats (Neo4j, FalkorDB, SVG), and enhancing confidence tagging for node relationships
- **Contributor pathway**: Standardized process in [`ARCHITECTURE.md`](https://github.com/Graphify-Labs/graphify/blob/main/ARCHITECTURE.md) enables community-driven language support without core architecture changes

## Frequently Asked Questions

### What bug fixes are currently prioritized in the Graphify roadmap?

The roadmap explicitly tracks five critical bugs (designated A through E) that affect the extraction pipeline's reliability. These are documented in the "roadmap bug-fixes" test case within [`tests/test_dart.py`](https://github.com/Graphify-Labs/graphify/blob/main/tests/test_dart.py) and target edge cases in AST traversal and edge inference during the extract phase.

### How can I add support for a new programming language to Graphify?

You can add support by implementing an `extract_<lang>` function in [`graphify/extract.py`](https://github.com/Graphify-Labs/graphify/blob/main/graphify/extract.py) that returns nodes and edges dictionaries, then registering the file extension in both [`graphify/detect.py`](https://github.com/Graphify-Labs/graphify/blob/main/graphify/detect.py) (updating `CODE_EXTENSIONS`) and [`graphify/watch.py`](https://github.com/Graphify-Labs/graphify/blob/main/graphify/watch.py) (updating `_WATCHED_EXTENSIONS`). Finally, add the tree-sitter grammar to [`pyproject.toml`](https://github.com/Graphify-Labs/graphify/blob/main/pyproject.toml) as an optional extra and provide tests in [`tests/test_languages.py`](https://github.com/Graphify-Labs/graphify/blob/main/tests/test_languages.py).

### Will Graphify support exporting to graph databases like Neo4j?

Yes, the development roadmap includes enhanced export formats specifically targeting **Neo4j** and **FalkorDB** integration, alongside SVG visualization outputs. These features are part of the analytics scalability theme aimed at improving integration with external graph analysis tools.

### Is the roadmap focused on stability or new features?

The roadmap balances both objectives simultaneously. The immediate focus is on **stabilization** through the five documented bug fixes in the extraction pipeline, while parallel efforts expand **language coverage** and **analytics capabilities**. This dual-track approach allows Graphify to harden existing functionality while growing its ecosystem.