How to Contribute to the Graphify Project: Architecture Guide and PR Workflow

Contribute to Graphify by submitting worked examples to the worked/ directory, reporting extraction bugs with cached outputs, or adding new language extractors via graphify/extract.py and detect.py, then opening a PR against the v8 branch.

Graphify transforms entire codebases into queryable local knowledge graphs, and the Graphify-Labs/graphify repository actively welcomes contributions across documentation, bug reports, and core functionality. Before submitting your first pull request, you must understand the seven-stage pipeline architecture that processes source files from detection through export.

Understanding the Graphify Architecture

The codebase follows a strict functional pipeline where each stage communicates via plain Python dictionaries and NetworkX graphs, guaranteeing a side-effect-free workflow that writes only under graphify-out/. According to ARCHITECTURE.md, the pipeline consists of seven distinct stages:

Contribution Pathways

The README.md outlines three primary ways to contribute to Graphify, each targeting different skill levels and interests.

Submitting Worked Examples

Run /graphify on a real-world corpus, save the output under worked/{slug}/, and write a thorough review.md documenting what the graph captured correctly and what it missed. Reference existing examples in the worked/ directory for formatting standards.

Reporting Extraction Bugs

Open an issue containing the problematic source file, the cached extraction from graphify-out/cache/, and a detailed description of which nodes or edges were missed. This helps improve the AST parsing logic in the extraction stage.

Adding New Language Extractors

Extend Graphify to support languages not yet parsed by implementing a new extractor module. This is the most technical contribution path and requires modifying multiple core files.

Step-by-Step: Adding a Language Extractor

To add support for a new language, follow the implementation steps enumerated in ARCHITECTURE.md:

  1. Implement the extraction function in graphify/extract.py (or a dedicated file) following the signature extract_<lang>(path: Path) -> dict. The function must return a dictionary with nodes and edges keys conforming to the JSON schema specified in ARCHITECTURE.md.

  2. Register the file suffix in the central dispatcher by adding an entry to the DISPATCH dictionary in extract().

  3. Update extension whitelists by adding the suffix to CODE_EXTENSIONS in graphify/detect.py and _WATCHED_EXTENSIONS in watch.py.

  4. Add tree-sitter dependency to pyproject.toml if the language requires a new parser.

  5. Create test fixtures under tests/fixtures/ and add a corresponding test case in tests/test_languages.py.

Example: Adding a Toy ".foo" Extractor

Create the extractor class in a new file:


# graphify/extractors/foo.py

from pathlib import Path
from .base import BaseExtractor

class FooExtractor(BaseExtractor):
    """Treats each line as a separate node."""

    def extract(self, path: Path) -> dict:
        nodes, edges = [], []
        for i, line in enumerate(path.read_text().splitlines(), start=1):
            node_id = f"{path}:{i}"
            nodes.append({
                "id": node_id,
                "label": line.strip(),
                "source_file": str(path),
                "source_location": f"L{i}"
            })
        return {"nodes": nodes, "edges": edges}

Register the extractor in the dispatcher:


# graphify/extract.py

from .extractors.foo import FooExtractor

DISPATCH = {
    ".py": extract_python,
    ".js": extract_javascript,
    ".foo": FooExtractor().extract,  # New entry

}

Update the detection whitelist:


# graphify/detect.py

CODE_EXTENSIONS = {".py", ".js", ".foo"}

Verify your implementation with a unit test:


# tests/test_languages.py

def test_foo_extractor(tmp_path):
    src = tmp_path / "example.foo"
    src.write_text("alpha\nbeta\n")
    result = graphify.extract.extract(str(src))
    assert len(result["nodes"]) == 2
    assert result["nodes"][0]["label"] == "alpha"

Local Development Setup

Before contributing, configure your local environment to run the full pipeline:


# Install the CLI tool

uv tool install graphifyy

# Or install from source for development

uv sync --all-extras

# Run the test suite to verify setup

uv run pytest -q

# Execute a local build

graphify .

# Output appears in graphify-out/ as graph.json, GRAPH_REPORT.md, and graph.html

Submitting Your Contribution

All pull requests must target the v8 branch. Follow this checklist before submission:

  • Branch from v8 and maintain the existing code style
  • Add unit tests for any new extractors or logic in tests/
  • Update documentation in README.md or ARCHITECTURE.md if you modify public APIs
  • For worked examples, include the graphify-out/ directory and review.md under worked/{slug}/
  • Ensure the full test suite passes with uv run pytest -q

Summary

  • Graphify uses a seven-stage pipeline (Detect, Extract, Build, Cluster, Analyze, Report, Export) that processes code into NetworkX graphs.
  • Contribute by adding worked examples to worked/, reporting extraction bugs with cache files, or implementing new language extractors.
  • New extractors require updating graphify/extract.py, graphify/detect.py, and tests/test_languages.py while conforming to the JSON schema in ARCHITECTURE.md.
  • Always target the v8 branch for pull requests and include comprehensive tests.

Frequently Asked Questions

Which branch should I target for pull requests?

Target the v8 branch for all contributions. The repository uses this as the main development branch, and maintainers will handle merging into release tags.

What data format must language extractors return?

Extractors must return a Python dictionary with nodes and edges keys following the JSON schema defined in ARCHITECTURE.md. Each node requires id, label, source_file, and source_location fields to integrate properly with the Build and Cluster stages.

How do I test my changes before submitting a PR?

Run the unit test suite with uv run pytest -q to verify logic changes. For extractor contributions, create a fixture in tests/fixtures/ and add a test case to tests/test_languages.py. For worked examples, verify that graphify . successfully generates graphify-out/graph.json and GRAPH_REPORT.md without errors.

Where does Graphify write its output files?

According to the architecture specification, the pipeline writes exclusively to the graphify-out/ directory, including graph.json (machine-readable), GRAPH_REPORT.md (human-readable), cache files, and HTML visualizations. The tool never modifies source files in-place.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →