Development Roadmap for Graphify: Pipeline Stability, Language Expansion, and Analytics

Graphify's development roadmap prioritizes stabilizing the core extraction pipeline through critical bug fixes, expanding language coverage via modular tree-sitter extractors, and scaling graph analytics with parallel clustering and enhanced export formats.

The development roadmap for Graphify (Graphify-Labs/graphify) focuses on hardening the code-to-graph extraction engine while broadening support for modern programming languages. According to the project's ARCHITECTURE.md and test suites, the maintainers are actively addressing five documented bugs in the extraction stages and standardizing the process for adding new language parsers. This roadmap reflects an open-ended, community-driven approach where contributors can prioritize new extractors or performance improvements based on demand.

Stabilizing the Core Extraction Pipeline

The immediate priority in the Graphify roadmap is production stability for the extraction pipeline. The architecture follows a strict four-stage flow defined in ARCHITECTURE.md: detect → extract → build → analyze.

Critical to this phase is the "roadmap bug-fixes" test case found in tests/test_dart.py, which explicitly tracks five known bugs (labeled A through E) affecting the extraction stages. These fixes target edge cases in AST parsing and edge inference that currently impact the reliability of the extract() dispatcher in graphify/extract.py. Resolving these issues ensures that the detect phase (handled by graphify/detect.py) and the extract phase produce consistent node and edge sets before the graph construction layer processes them.

Expanding Language Coverage and Extractor Quality

The second pillar of the roadmap emphasizes language ecosystem growth. Graphify currently ships with 36 tree-sitter grammars and provides a documented, repeatable process in ARCHITECTURE.md for adding support for additional languages.

To implement a new language extractor, contributors must:

Target languages explicitly mentioned for future inclusion include Rust and Kotlin, though the modular architecture supports any language with a tree-sitter grammar. The optional extras system allows users to install language-specific parsers on demand without bloating the core installation.

Improving Graph Analytics and Scalability

The third strategic focus addresses performance and export capabilities for large-scale codebases. Currently, community detection and "god-node" identification rely on Leiden clustering algorithms (referenced in benchmark.py) and LLM-free heuristics with three confidence levels: EXTRACTED, INFERRED, and AMBIGUOUS.

Future milestones include:

  • Parallel processing: The graphify cluster-only command already accepts --max-concurrency and --resolution flags, indicating active development toward faster, multi-threaded community detection
  • Intelligent deduplication: Automatic merging of duplicate nodes with improved confidence tagging
  • Database integration: Native export formats for Neo4j, FalkorDB, and SVG visualization
  • MCP server support: Real-time querying integration with Model Context Protocol servers for AI-assisted code analysis

How to Contribute: Adding a New Language Extractor

The following example demonstrates the roadmap's standardized process for adding a "FooLang" extractor, following the implementation pattern established in graphify/extract.py:


# graphify/extract.py

from pathlib import Path
import tree_sitter  # ← the parser for the new language

def extract_foolang(path: Path) -> dict:
    """Parse a FooLang file and return a nodes/edges dict."""
    source = path.read_text()
    parser = tree_sitter.Parser()
    parser.set_language(tree_sitter.Language('build/my-languages.so', 'foolang'))
    tree = parser.parse(bytes(source, "utf8"))

    # Walk the AST, collect nodes and explicit edges.

    nodes, edges = [], []
    # ... (AST walking logic omitted for brevity) ...

    # Second-pass call-graph to infer additional `calls` edges.

    # (See existing extractors for reference.)

    return {"nodes": nodes, "edges": edges}

# Dispatch logic (already present in the file):

def extract(path: Path) -> dict:
    if path.suffix == ".foo":
        return extract_foolang(path)
    # existing language dispatches …

To complete the integration, update CODE_EXTENSIONS in graphify/detect.py and _WATCHED_EXTENSIONS in graphify/watch.py to include .foo, then add the grammar dependency to pyproject.toml.

Summary

  • Immediate priorities: Resolving bugs A-E in the extraction pipeline, documented in tests/test_dart.py, to stabilize the detect-extract-build-analyze flow
  • Language expansion: Supporting additional languages like Rust and Kotlin through the modular extract_<lang> pattern and optional dependency management
  • Scalability improvements: Implementing parallel community detection (--max-concurrency), expanding export formats (Neo4j, FalkorDB, SVG), and enhancing confidence tagging for node relationships
  • Contributor pathway: Standardized process in ARCHITECTURE.md enables community-driven language support without core architecture changes

Frequently Asked Questions

What bug fixes are currently prioritized in the Graphify roadmap?

The roadmap explicitly tracks five critical bugs (designated A through E) that affect the extraction pipeline's reliability. These are documented in the "roadmap bug-fixes" test case within tests/test_dart.py and target edge cases in AST traversal and edge inference during the extract phase.

How can I add support for a new programming language to Graphify?

You can add support by implementing an extract_<lang> function in graphify/extract.py that returns nodes and edges dictionaries, then registering the file extension in both graphify/detect.py (updating CODE_EXTENSIONS) and graphify/watch.py (updating _WATCHED_EXTENSIONS). Finally, add the tree-sitter grammar to pyproject.toml as an optional extra and provide tests in tests/test_languages.py.

Will Graphify support exporting to graph databases like Neo4j?

Yes, the development roadmap includes enhanced export formats specifically targeting Neo4j and FalkorDB integration, alongside SVG visualization outputs. These features are part of the analytics scalability theme aimed at improving integration with external graph analysis tools.

Is the roadmap focused on stability or new features?

The roadmap balances both objectives simultaneously. The immediate focus is on stabilization through the five documented bug fixes in the extraction pipeline, while parallel efforts expand language coverage and analytics capabilities. This dual-track approach allows Graphify to harden existing functionality while growing its ecosystem.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →