How to Contribute to the Graphify Project: Architecture Guide and PR Workflow
Contribute to Graphify by submitting worked examples to the worked/ directory, reporting extraction bugs with cached outputs, or adding new language extractors via graphify/extract.py and detect.py, then opening a PR against the v8 branch.
Graphify transforms entire codebases into queryable local knowledge graphs, and the Graphify-Labs/graphify repository actively welcomes contributions across documentation, bug reports, and core functionality. Before submitting your first pull request, you must understand the seven-stage pipeline architecture that processes source files from detection through export.
Understanding the Graphify Architecture
The codebase follows a strict functional pipeline where each stage communicates via plain Python dictionaries and NetworkX graphs, guaranteeing a side-effect-free workflow that writes only under graphify-out/. According to ARCHITECTURE.md, the pipeline consists of seven distinct stages:
- Detect (
graphify/detect.py): Walks the file tree and returns a filtered list of paths based onCODE_EXTENSIONS. - Extract (
graphify/extract.py): Parses files using tree-sitter ASTs via language-specific extractors to collect nodes and edges. - Build (
graphify/build.py): Assembles extraction dictionaries into a NetworkX graph structure. - Cluster (
graphify/cluster.py): Runs Leiden community detection and labels each node with acommunityattribute. - Analyze (
graphify/analyze.py): Computes "god nodes", surprising connections, and suggested questions. - Report (
graphify/report.py): Renders a human-readableGRAPH_REPORT.mdwith highlights and token usage statistics. - Export (
graphify/export.py): Producesgraph.json,graph.html, Obsidian vaults, and SVG visualizations.
Contribution Pathways
The README.md outlines three primary ways to contribute to Graphify, each targeting different skill levels and interests.
Submitting Worked Examples
Run /graphify on a real-world corpus, save the output under worked/{slug}/, and write a thorough review.md documenting what the graph captured correctly and what it missed. Reference existing examples in the worked/ directory for formatting standards.
Reporting Extraction Bugs
Open an issue containing the problematic source file, the cached extraction from graphify-out/cache/, and a detailed description of which nodes or edges were missed. This helps improve the AST parsing logic in the extraction stage.
Adding New Language Extractors
Extend Graphify to support languages not yet parsed by implementing a new extractor module. This is the most technical contribution path and requires modifying multiple core files.
Step-by-Step: Adding a Language Extractor
To add support for a new language, follow the implementation steps enumerated in ARCHITECTURE.md:
-
Implement the extraction function in
graphify/extract.py(or a dedicated file) following the signatureextract_<lang>(path: Path) -> dict. The function must return a dictionary withnodesandedgeskeys conforming to the JSON schema specified inARCHITECTURE.md. -
Register the file suffix in the central dispatcher by adding an entry to the
DISPATCHdictionary inextract(). -
Update extension whitelists by adding the suffix to
CODE_EXTENSIONSingraphify/detect.pyand_WATCHED_EXTENSIONSinwatch.py. -
Add tree-sitter dependency to
pyproject.tomlif the language requires a new parser. -
Create test fixtures under
tests/fixtures/and add a corresponding test case intests/test_languages.py.
Example: Adding a Toy ".foo" Extractor
Create the extractor class in a new file:
# graphify/extractors/foo.py
from pathlib import Path
from .base import BaseExtractor
class FooExtractor(BaseExtractor):
"""Treats each line as a separate node."""
def extract(self, path: Path) -> dict:
nodes, edges = [], []
for i, line in enumerate(path.read_text().splitlines(), start=1):
node_id = f"{path}:{i}"
nodes.append({
"id": node_id,
"label": line.strip(),
"source_file": str(path),
"source_location": f"L{i}"
})
return {"nodes": nodes, "edges": edges}
Register the extractor in the dispatcher:
# graphify/extract.py
from .extractors.foo import FooExtractor
DISPATCH = {
".py": extract_python,
".js": extract_javascript,
".foo": FooExtractor().extract, # New entry
}
Update the detection whitelist:
# graphify/detect.py
CODE_EXTENSIONS = {".py", ".js", ".foo"}
Verify your implementation with a unit test:
# tests/test_languages.py
def test_foo_extractor(tmp_path):
src = tmp_path / "example.foo"
src.write_text("alpha\nbeta\n")
result = graphify.extract.extract(str(src))
assert len(result["nodes"]) == 2
assert result["nodes"][0]["label"] == "alpha"
Local Development Setup
Before contributing, configure your local environment to run the full pipeline:
# Install the CLI tool
uv tool install graphifyy
# Or install from source for development
uv sync --all-extras
# Run the test suite to verify setup
uv run pytest -q
# Execute a local build
graphify .
# Output appears in graphify-out/ as graph.json, GRAPH_REPORT.md, and graph.html
Submitting Your Contribution
All pull requests must target the v8 branch. Follow this checklist before submission:
- Branch from
v8and maintain the existing code style - Add unit tests for any new extractors or logic in
tests/ - Update documentation in
README.mdorARCHITECTURE.mdif you modify public APIs - For worked examples, include the
graphify-out/directory andreview.mdunderworked/{slug}/ - Ensure the full test suite passes with
uv run pytest -q
Summary
- Graphify uses a seven-stage pipeline (Detect, Extract, Build, Cluster, Analyze, Report, Export) that processes code into NetworkX graphs.
- Contribute by adding worked examples to
worked/, reporting extraction bugs with cache files, or implementing new language extractors. - New extractors require updating
graphify/extract.py,graphify/detect.py, andtests/test_languages.pywhile conforming to the JSON schema inARCHITECTURE.md. - Always target the
v8branch for pull requests and include comprehensive tests.
Frequently Asked Questions
Which branch should I target for pull requests?
Target the v8 branch for all contributions. The repository uses this as the main development branch, and maintainers will handle merging into release tags.
What data format must language extractors return?
Extractors must return a Python dictionary with nodes and edges keys following the JSON schema defined in ARCHITECTURE.md. Each node requires id, label, source_file, and source_location fields to integrate properly with the Build and Cluster stages.
How do I test my changes before submitting a PR?
Run the unit test suite with uv run pytest -q to verify logic changes. For extractor contributions, create a fixture in tests/fixtures/ and add a test case to tests/test_languages.py. For worked examples, verify that graphify . successfully generates graphify-out/graph.json and GRAPH_REPORT.md without errors.
Where does Graphify write its output files?
According to the architecture specification, the pipeline writes exclusively to the graphify-out/ directory, including graph.json (machine-readable), GRAPH_REPORT.md (human-readable), cache files, and HTML visualizations. The tool never modifies source files in-place.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →