Graphify's Global Cross-Project Graph: Unified Multi-Repository Knowledge Base

Graphify's global cross-project graph is a unified knowledge base stored at ~/.graphify/global.json that merges individual project graphs using prefixed IDs and deduplication to enable multi-repository analysis.

Graphify builds individual knowledge graphs for each codebase it processes. To support analysis across multiple repositories, Graphify implements a global cross-project graph that aggregates these individual graphs into a single queryable view. This feature, available in the Graphify-Labs/graphify repository, allows developers to trace relationships and dependencies across their entire codebase ecosystem.

What is the Global Cross-Project Graph?

The global cross-project graph is a merged representation of multiple project-level knowledge graphs. When you process a codebase with Graphify, it creates a standalone graph for that project. The global graph collects these individual graphs into a single JSON file, enabling queries that span across repository boundaries. This unified view preserves the same schema used for single-project graphs—defined in ARCHITECTURE.md—so all downstream analysis tools operate unchanged on the global view.

How the Global Graph Works

Registration and Storage

Projects enter the global graph through the --global flag during extraction or via manual registration. When you run graphify extract ./my-project --global --as myproject, the CLI automatically registers the freshly built graph using the graphify global add command. All registered graphs persist in a JSON file at ~/.graphify/global.json. You can locate this file at any time by running graphify global path.

Collision-Free ID Management

To prevent namespace collisions when merging graphs from different sources, Graphify automatically prefixes every node ID with <repo>:: (the repository name or tag). This prefixing ensures that a function named parse_data in project A remains distinct from parse_data in project B, avoiding silent overwrites while maintaining relationships within each original project.

Deduplication Logic

The global graph implements intelligent deduplication for external dependencies. When two projects reference the same third-party library—such as both importing requests—Graphify creates a single canonical node for that dependency and rewires all edges to point to this unified representation. This deduplication reduces graph size and accurately reflects shared dependencies across your codebase ecosystem.

Incremental Update Mechanism

Adding a project to the global graph multiple times is computationally cheap. Graphify computes a hash of the incoming graph and compares it against already-registered versions. If the hash matches an existing entry, the system skips re-ingestion entirely. This hashing mechanism keeps the global graph up-to-date without unnecessary processing overhead.

Working with the Global Graph

You can interact with the global graph through both the CLI and Python API.

CLI Commands:


# Extract and automatically register a project

graphify extract ./my-project --global --as myproject

# Manually add an existing graph file

graphify global add graphify-out/graph.json --as myproject

# List all registered projects with node/edge counts

graphify global list

# Remove a project from the global graph

graphify global remove myproject

# Display the path to global.json

graphify global path

Python API:

For programmatic access, import the global_graph module and load the unified graph:

from graphify import global_graph

# Load the current global graph

g = global_graph.load_global()

# Find all projects that import the same library

libs = [n for n, data in g.nodes(data=True) if "requests" in data.get("label", "")]
print(f"Projects using `requests`: {set(n.split('::')[0] for n in libs)}")

Core Implementation Files

The global graph functionality spans several key files in the Graphify codebase:

  • graphify/cli.py: Implements the global add, global remove, global list, and global path subcommands that manage the registry of projects.
  • graphify/global_graph.py: Handles the loading, merging, deduplication logic, and persistence operations for the ~/.graphify/global.json file.
  • graphify/export.py: Generates the project-level graphify-out/graph.json files that serve as input to the global graph system.
  • ARCHITECTURE.md: Documents the extraction schema and overall pipeline (detect → extract → build → cluster → analyze → report → export) used by both project-level and global graphs.

Summary

  • Storage Location: The global graph resides at ~/.graphify/global.json, accessible via graphify global path.
  • ID Prefixing: Node IDs use <repo>:: prefixes to prevent collisions across projects.
  • Deduplication: External dependencies automatically merge into canonical nodes, creating accurate cross-project dependency maps.
  • Incremental Updates: Hash-based comparison prevents redundant re-ingestion of unchanged projects.
  • Unified Schema: The global graph uses the same schema as individual projects, ensuring compatibility with existing analysis tools.

Frequently Asked Questions

Where is the global cross-project graph stored?

Graphify stores the global graph at ~/.graphify/global.json on your local filesystem. You can verify the exact path by running the command graphify global path in your terminal. This JSON file contains the merged nodes and edges from all registered projects.

How does Graphify prevent ID collisions between different projects?

Graphify automatically prefixes every node ID with <repo>:: (using the repository name or tag you specified during registration). This namespace prefixing ensures that identical identifiers from different projects remain distinct in the merged graph, preventing silent data overwrites.

Can I incrementally update the global graph without rebuilding everything?

Yes. Graphify computes a hash of each incoming project graph and compares it against existing entries. If the hash matches a previously registered version, the system skips re-ingestion. This makes it efficient to run graphify extract --global repeatedly as your codebase evolves.

How do I query dependencies across multiple repositories?

Use the Python API to load the global graph with global_graph.load_global(), then traverse nodes using standard graph methods. Filter nodes by their label attribute and extract the repository prefix from the node ID (splitting on ::) to identify which projects share specific dependencies or call relationships.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →