How Graphify's Global Graph Feature Works for Multi-Repo Queries
Graphify enables multi-repo queries by maintaining a unified knowledge graph under ~/.graphify that namespaces each repository with unique prefixes, deduplicates external library references, and persists the merged result to a global JSON file.
Graphify is a Python tool that constructs knowledge graphs from codebases to map symbol relationships. When you need to analyze dependencies and connections across multiple projects, Graphify's global graph feature aggregates individual repository graphs into a single queryable structure, allowing you to perform cross-repository analysis using standard graph traversal techniques.
The Architecture of the Global Graph
The global graph lives under ~/.graphify/global-graph.json and serves as a unified namespace for all added repositories. When you add a repository using global_add, Graphify executes a seven-step process to ensure clean integration without data collisions.
Namespace Isolation via Prefixing
Every node ID in the repository gets prefixed with repo_tag:: to prevent collisions when different projects define symbols with identical names. This transformation happens in prefix_graph_for_global located in graphify/build.py (lines 93-100). The function creates a copy of the graph, rewrites node IDs with the repository-specific prefix, and adds two critical attributes: repo indicating the source repository, and local_id preserving the original identifier.
External Library Deduplication
Nodes lacking a source_file attribute represent external libraries or third-party dependencies. Rather than creating duplicate entries for common libraries used across multiple repositories, Graphify merges these nodes by their label attribute. This deduplication logic resides in global_add within graphify/global_graph.py (lines 21-33), ensuring that if two repositories reference the same external symbol, the global graph maintains a single canonical node.
Manifest and Persistence
Graphify maintains a _GLOBAL_MANIFEST that tracks metadata for each added repository, including the source path, node count, timestamp, and a SHA-256 hash of the original file. The functions _load_manifest and _save_manifest (lines 15-45 of global_graph.py) handle safe read/write operations with automatic backup generation if parsing errors occur. Before merging new data, Graphify prunes stale nodes belonging to the same repo_tag using prune_repo_from_graph in graphify/build.py (lines 9-13).
Working with the Global Graph
Adding Repositories to the Global Graph
To incorporate a repository into the global namespace, call global_add with the path to the local graph JSON and a unique repository tag:
from pathlib import Path
from graphify.global_graph import global_add, global_list
repo_path = Path("/my/project/graph.json") # output of `graphify build`
tag = "myproj" # unique identifier for the repo
result = global_add(repo_path, tag)
print(result) # → {'repo_tag': 'myproj', 'nodes_added': 123, 'nodes_removed': 0, 'skipped': False}
The global_add function (lines 77-86 of global_graph.py) handles the complete workflow: loading the local graph, prefixing node IDs, pruning existing data for that tag, deduplicating externals, and merging into the global structure.
Querying Across Multiple Repositories
Once aggregated, you can load the global graph and run queries that span repository boundaries using NetworkX:
import json
import networkx as nx
from graphify.global_graph import global_path
from networkx.readwrite import json_graph as jg
with open(global_path(), "r", encoding="utf-8") as f:
data = json.load(f)
G = jg.node_link_graph(data, edges="links")
# Example: find all nodes that depend on a symbol named "User"
dependents = [n for n, d in G.nodes(data=True) if d.get("label") == "User"]
# `dependents` now contains nodes from any repo that imported `User`.
Because each node carries a repo attribute, you can scope queries to specific projects or analyze the complete union of all repositories.
Repository Maintenance
Remove obsolete repositories or inspect the current global state using the following utilities:
from graphify.global_graph import global_list, global_remove
# List all repositories in the global graph
print(global_list())
# {'myproj': {'added_at': '2024‑06‑15T12:34:56Z', 'source_path': '/my/project/graph.json', …}}
# Remove a specific repository
removed = global_remove("myproj")
print(f"Removed {removed} nodes belonging to 'myproj'")
Key Implementation Files
graphify/global_graph.py: Contains the core API includingglobal_add,global_remove,global_list, and manifest management functions.graphify/build.py: Implements helper functionsprefix_graph_for_globalfor ID namespacing andprune_repo_from_graphfor cleaning stale data.tests/test_global_graph.py: Comprehensive test suite validating the global graph workflow including collision handling, deduplication, and removal operations.
Summary
- Graphify stores multi-repo data in
~/.graphify/global-graph.jsonas a unified NetworkX-compatible graph. - Node collision prevention is achieved by prefixing all IDs with
repo_tag::before merging. - External libraries are deduplicated by merging nodes without
source_fileattributes based on theirlabel. - Repository metadata is tracked in a
_GLOBAL_MANIFESTwith safe persistence and hash verification. - Cross-repo queries leverage the
repoattribute on nodes to filter or combine data from specific projects.
Frequently Asked Questions
Where is the global graph stored on disk?
The global graph persists to ~/.graphify/global-graph.json, while repository metadata lives in a manifest file within the same directory. These paths are managed internally by the global_path() utility and _GLOBAL_MANIFEST constants in graphify/global_graph.py.
How does Graphify prevent node ID collisions between repositories?
Graphify prevents collisions by prefixing every node ID with a repository-specific tag (e.g., myproj::) using the prefix_graph_for_global function in graphify/build.py. This ensures that identical symbol names from different repositories remain distinct in the global namespace.
What happens to external library references when merging multiple repositories?
External nodes—those without a source_file attribute—are automatically deduplicated by their label attribute during the merge process. This means two repositories referencing the same third-party library will share a single node in the global graph, reducing redundancy while maintaining accurate relationship mapping.
Can I query specific repositories within the global graph?
Yes. Every node in the global graph includes a repo attribute indicating its source repository. You can filter NetworkX queries using this attribute to analyze single projects, or omit the filter to query across all added repositories simultaneously.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →