Graphify's Global Cross-Project Graph: Unified Multi-Repository Knowledge Base
Graphify's global cross-project graph is a unified knowledge base stored at ~/.graphify/global.json that merges individual project graphs using prefixed IDs and deduplication to enable multi-repository analysis.
Graphify builds individual knowledge graphs for each codebase it processes. To support analysis across multiple repositories, Graphify implements a global cross-project graph that aggregates these individual graphs into a single queryable view. This feature, available in the Graphify-Labs/graphify repository, allows developers to trace relationships and dependencies across their entire codebase ecosystem.
What is the Global Cross-Project Graph?
The global cross-project graph is a merged representation of multiple project-level knowledge graphs. When you process a codebase with Graphify, it creates a standalone graph for that project. The global graph collects these individual graphs into a single JSON file, enabling queries that span across repository boundaries. This unified view preserves the same schema used for single-project graphs—defined in ARCHITECTURE.md—so all downstream analysis tools operate unchanged on the global view.
How the Global Graph Works
Registration and Storage
Projects enter the global graph through the --global flag during extraction or via manual registration. When you run graphify extract ./my-project --global --as myproject, the CLI automatically registers the freshly built graph using the graphify global add command. All registered graphs persist in a JSON file at ~/.graphify/global.json. You can locate this file at any time by running graphify global path.
Collision-Free ID Management
To prevent namespace collisions when merging graphs from different sources, Graphify automatically prefixes every node ID with <repo>:: (the repository name or tag). This prefixing ensures that a function named parse_data in project A remains distinct from parse_data in project B, avoiding silent overwrites while maintaining relationships within each original project.
Deduplication Logic
The global graph implements intelligent deduplication for external dependencies. When two projects reference the same third-party library—such as both importing requests—Graphify creates a single canonical node for that dependency and rewires all edges to point to this unified representation. This deduplication reduces graph size and accurately reflects shared dependencies across your codebase ecosystem.
Incremental Update Mechanism
Adding a project to the global graph multiple times is computationally cheap. Graphify computes a hash of the incoming graph and compares it against already-registered versions. If the hash matches an existing entry, the system skips re-ingestion entirely. This hashing mechanism keeps the global graph up-to-date without unnecessary processing overhead.
Working with the Global Graph
You can interact with the global graph through both the CLI and Python API.
CLI Commands:
# Extract and automatically register a project
graphify extract ./my-project --global --as myproject
# Manually add an existing graph file
graphify global add graphify-out/graph.json --as myproject
# List all registered projects with node/edge counts
graphify global list
# Remove a project from the global graph
graphify global remove myproject
# Display the path to global.json
graphify global path
Python API:
For programmatic access, import the global_graph module and load the unified graph:
from graphify import global_graph
# Load the current global graph
g = global_graph.load_global()
# Find all projects that import the same library
libs = [n for n, data in g.nodes(data=True) if "requests" in data.get("label", "")]
print(f"Projects using `requests`: {set(n.split('::')[0] for n in libs)}")
Core Implementation Files
The global graph functionality spans several key files in the Graphify codebase:
graphify/cli.py: Implements theglobal add,global remove,global list, andglobal pathsubcommands that manage the registry of projects.graphify/global_graph.py: Handles the loading, merging, deduplication logic, and persistence operations for the~/.graphify/global.jsonfile.graphify/export.py: Generates the project-levelgraphify-out/graph.jsonfiles that serve as input to the global graph system.ARCHITECTURE.md: Documents the extraction schema and overall pipeline (detect → extract → build → cluster → analyze → report → export) used by both project-level and global graphs.
Summary
- Storage Location: The global graph resides at
~/.graphify/global.json, accessible viagraphify global path. - ID Prefixing: Node IDs use
<repo>::prefixes to prevent collisions across projects. - Deduplication: External dependencies automatically merge into canonical nodes, creating accurate cross-project dependency maps.
- Incremental Updates: Hash-based comparison prevents redundant re-ingestion of unchanged projects.
- Unified Schema: The global graph uses the same schema as individual projects, ensuring compatibility with existing analysis tools.
Frequently Asked Questions
Where is the global cross-project graph stored?
Graphify stores the global graph at ~/.graphify/global.json on your local filesystem. You can verify the exact path by running the command graphify global path in your terminal. This JSON file contains the merged nodes and edges from all registered projects.
How does Graphify prevent ID collisions between different projects?
Graphify automatically prefixes every node ID with <repo>:: (using the repository name or tag you specified during registration). This namespace prefixing ensures that identical identifiers from different projects remain distinct in the merged graph, preventing silent data overwrites.
Can I incrementally update the global graph without rebuilding everything?
Yes. Graphify computes a hash of each incoming project graph and compares it against existing entries. If the hash matches a previously registered version, the system skips re-ingestion. This makes it efficient to run graphify extract --global repeatedly as your codebase evolves.
How do I query dependencies across multiple repositories?
Use the Python API to load the global graph with global_graph.load_global(), then traverse nodes using standard graph methods. Filter nodes by their label attribute and extract the repository prefix from the node ID (splitting on ::) to identify which projects share specific dependencies or call relationships.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →