Core Components of Graphify: Local-First Knowledge Graph Architecture Explained

Graphify transforms entire projects—code, documents, PDFs, images, and videos—into structured, queryable knowledge graphs using a three-pass extraction pipeline implemented in tightly-coupled Python modules.

Graphify is an open-source, local-first knowledge graph engine developed by Graphify-Labs. Unlike traditional code search tools that rely on grepping, Graphify constructs a rich, interconnected graph of your entire project that you can query programmatically or via natural language. Understanding the core components of Graphify reveals how it maintains complete data privacy while delivering AI-powered insights through its modular Python architecture.

Three-Pass Extraction Pipeline

The foundation of Graphify’s architecture is a sequential three-pass pipeline that isolates heavyweight LLM operations to minimize token costs while keeping most extractions offline.

Pass 1 – Code Structure Extraction (Offline)

The first pass utilizes tree-sitter parsers to analyze every supported source file without requiring API calls. In graphify/extract.py, the system walks the file tree and dispatches to 36 language-specific extractors located under graphify/extractors/ (including python.py, typescript.py, and rust.py). These modules extract classes, functions, imports, call graphs, and inline comments, producing a raw code-only graph with precise AST-based relationships.

Pass 2 – Audio and Video Transcription (Local)

For multimedia content, graphify/transcribe.py leverages faster-whisper to process audio and video files entirely locally. The transcription engine seeds results with existing "god nodes" to focus extraction on domain-relevant content, ensuring technical vocabulary is preserved without sending media to external services.

Pass 3 – Semantic Document Processing (LLM-Driven)

Non-code artifacts—PDFs, images, and documentation—undergo parallel processing in graphify/llm.py. This module coordinates sub-agents that interface with configured backends (Claude, OpenAI, or Ollama) using the system prompt defined in references/extraction-spec.md. Each artifact is transformed into JSON nodes and edges, with confidence scores assigned based on extraction certainty.

Build Orchestration and Merging

The build orchestrator in graphify/build.py wires all three passes together. It executes the pipeline sequentially, merges partial graphs from each phase, handles caching via graphify/cache.py, and triggers post-processing steps including community detection and export generation.

Core Python Modules

Beyond the pipeline itself, Graphify consists of specialized modules that handle distinct responsibilities:

CLI and Entry Point – graphify/cli.py parses all graphify commands (graphify ., graphify query, graphify serve), serving as the user-facing interface to the engine.

Symbol Resolution – graphify/symbol_resolution.py resolves cross-file references for calls, imports, inherits, and mixes_in relationships. When AST analysis alone is insufficient, it creates INFERRED edges with confidence scores to maintain graph connectivity.

Community Detection – graphify/cluster.py applies the Leiden algorithm to group nodes into communities, automatically labeling clusters via targeted LLM calls to provide semantic groupings of related code and documentation.

MCP and HTTP Server – graphify/serve.py exposes the graph via a streaming RPC endpoint supporting methods like query_graph, shortest_path, and get_node. This Model Context Protocol (MCP) server enables IDE assistants and AI agents to query the graph over HTTP or stdio transports.

Hooks and Skill Installation – graphify/install.py generates platform-specific integration files including AGENTS.md and .cursor/rules/, allowing AI assistants to automatically query the Graphify knowledge base before performing operations.

Cache Management – graphify/cache.py implements SHA-256 fingerprinting for both AST and semantic extracts, ensuring unchanged files skip reprocessing in subsequent builds.

Exporters – graphify/export.py and sub-modules under graphify/exporters/ convert the internal NetworkX graph into multiple formats: HTML visualizations, Neo4j Cypher scripts, GraphDB files, and SVG diagrams.

Validation and Security – graphify/validate.py performs consistency checks on generated graphs, while graphify/security.py enforces safe handling of secrets and sensitive data throughout the extraction process.

Reporting – graphify/report.py generates human-readable outputs including GRAPH_REPORT.md and interactive graph.html files for manual exploration.

Data Model and Graph Structure

The final output is a NetworkX node-link JSON (graph.json) containing typed nodes and weighted edges:

  • Nodes include fields for id, label, file_type (code, document, image), source_file, and optional summary text.
  • Edges specify source, target, relation (e.g., calls, imports, semantically_similar_to), confidence level (EXTRACTED, INFERRED, AMBIGUOUS), and numeric confidence_score for inferred relationships.
  • Hyperedges are stored under G.graph["hyperedges"] to represent group relationships such as "module contains these functions."

This schema is formally documented in docs/how-it-works.md and docs/node-summaries-rfc.md.

Practical Implementation Examples

Building a Graph from a Local Repository

Install the CLI and run the extractor:


# Install via uv or pipx

uv tool install graphifyy

# Extract current directory

graphify .

This executes graphify/build.py, which internally calls graphify/extract.py (Pass 1), graphify/transcribe.py (Pass 2 when media present), and graphify/llm.py (Pass 3), outputting to graphify-out/graph.json and graphify-out/graph.html.

Querying the Graph Programmatically

from graphify import serve, querylog

# Load the generated graph

graph = serve.load_graph("graphify-out/graph.json")

# Find shortest path between concepts

path = serve.shortest_path(graph, "FastAPI", "ModelField")
print(" → ".join(path))

# Natural language query using configured LLM

answer = serve.query_graph(
    graph,
    "What connects authentication to the database?",
    backend="openai"  # or "claude", "ollama"

)
print(answer)

Running the MCP Server

Start the HTTP server for IDE integration:

python -m graphify.serve graphify-out/graph.json --transport http --port 8080

Clients can now invoke query_graph, get_node, or shortest_path via HTTP requests or stdio streams.

Summary

  • Graphify is a local-first knowledge graph engine that eliminates the need for grepping by structuring entire projects into queryable graphs.
  • The three-pass pipeline isolates LLM costs by processing code offline (Pass 1), media locally (Pass 2), and documents via AI only when necessary (Pass 3).
  • Core modules include graphify/build.py (orchestration), graphify/extract.py (AST parsing), graphify/llm.py (semantic extraction), and graphify/serve.py (MCP server).
  • The system generates NetworkX JSON graphs with typed nodes, confidence-weighted edges, and hyperedges for complex relationships.
  • Integration hooks in graphify/install.py enable seamless assistant workflows for Claude, Cursor, and other AI agents.

Frequently Asked Questions

What file types does Graphify support for extraction?

Graphify supports code files across 36 languages via tree-sitter (including Python, TypeScript, Rust, and Java), audio/video formats via faster-whisper transcription, and document formats including PDFs and images via LLM-based extraction in graphify/llm.py.

How does Graphify handle cross-file references in code?

The graphify/symbol_resolution.py module resolves cross-file references such as function calls, imports, and inheritance chains. When static analysis cannot definitively resolve a symbol, it creates INFERRED edges marked with confidence scores rather than omitting the relationship entirely.

Can Graphify work entirely offline?

Yes. Pass 1 (code extraction) and Pass 2 (audio/video transcription) run completely offline using local parsers and faster-whisper. Only Pass 3 requires internet connectivity when processing documents and images through external LLM providers like Claude or OpenAI, though this can be configured to use local Ollama instances for fully air-gapped operation.

What is the output format of the generated knowledge graph?

Graphify outputs a standard NetworkX node-link JSON file (graph.json) containing nodes with metadata (type, source file, summaries) and edges with relationship types and confidence levels. The system also generates interactive HTML reports (graph.html) and can export to graph databases like Neo4j via graphify/export.py.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →