How Multi-Repo Intelligence Works with CROSS_* Edges in codebase-memory-mcp
The codebase-memory-mcp server creates bidirectional CROSS_* edges between repositories by matching HTTP routes, async topics, and channels across projects during a dedicated indexing pass, enabling cross-repo queries with sub-millisecond latency.
The codebase-memory-mcp repository implements a sophisticated multi-repo intelligence system that unifies independent codebases into a single queryable knowledge graph. By leveraging specialized CROSS_* edge types, the server enables AI agents to trace dependencies and communication patterns across repository boundaries. This architecture allows tools like Claude Code and Codex CLI to answer structural questions such as "which services call this HTTP endpoint in any other repo?" with the same performance as intra-repo queries.
The Cross-Repo Discovery Pipeline
Discovery and Canonicalization
The pipeline begins in src/pipeline/pass_cross_repo.c, where the system scans every HTTP_CALLS, ASYNC_CALLS, or EMITS edge in a source project. For each edge, it searches all target projects for matching Route or Channel nodes. To handle service-specific hostnames, the cr_url_path function strips URL host components, canonicalizing paths like /v2/x for comparison regardless of deployment environment.
Bidirectional Edge Insertion
When matches are detected, the system creates bidirectional edges using the insert_cross_edge function. These edges are inserted into both the source and target SQLite databases, ensuring each repository maintains a complete local view of its external relationships. The edges are idempotent—unique on the composite key (source_id, target_id, type)—preventing duplication during repeated indexing passes.
Storage Architecture and Edge Properties
All CROSS_* edges reside in the standard edges table alongside intra-repo relationships. The build_cross_props function constructs a JSON payload containing the target project name, function, file path, and protocol details (such as HTTP method or channel identifier). Each repository maintains its own SQLite file (<project>.db), allowing the pass to open both source and target databases simultaneously, insert the edge into each, and close them without requiring a central server.
Supported Cross-Repository Edge Types
According to the source code in src/mcp/mcp.c (lines 2818-2822), the engine supports six distinct cross-repo edge types:
CROSS_HTTP_CALLS– Links HTTP client calls to server route handlersCROSS_ASYNC_CALLS– Connects async function invocations across servicesCROSS_CHANNEL– Maps pub/sub channel emissions to listenersCROSS_GRPC_CALLS– Tracks gRPC method calls between repositoriesCROSS_GRAPHQL_CALLS– Identifies GraphQL query/mutation relationshipsCROSS_TRPC_CALLS– Links tRPC procedure calls across projects
Querying Cross-Repo Links
The MCP API exposes search_graph, trace_path, and get_architecture tools that automatically include CROSS_* edges when processing queries. For example, to find all services calling a specific endpoint across repositories:
{
"tool": "query_graph",
"params": {
"project": "my-project",
"query": "MATCH (src)-[:CROSS_HTTP_CALLS]->(tgt) WHERE src.name = 'fetchUser' RETURN tgt.project, tgt.name, tgt.properties_json"
}
}
The response includes the target_project field from the edge payload, enabling agents to construct complete multi-repo call chains.
Integration with AI Agents and CLI Tools
During indexing, the install command automatically enables the cross-repo pass for each project. Before recomputing links, the system removes existing CROSS_* edges to prevent stale connections, as implemented in src/pipeline/pass_cross_repo.c (lines 11-18).
Manual Edge Insertion
While typically automated, you can manually insert cross-repo edges using the C API:
// Assume store points to the source project's SQLite store
const char *props = NULL;
char buf[2048];
build_cross_props(buf, sizeof(buf),
"target-repo", // target project name
"handlerFunc", // target function name
"src/file.go", // target file
"/api/v1/items", // URL or channel identifier
"method", "GET"); // optional extra key/value
props = buf;
insert_cross_edge(store, "source-repo", caller_id, handler_id,
"CROSS_HTTP_CALLS", props);
CLI Query Examples
Retrieve all cross-repo HTTP edges for a project:
codebase-memory-mcp cli query_graph '{
"project":"my-repo",
"query":"MATCH (s)-[e:CROSS_HTTP_CALLS]->(t) RETURN s.name, e.type, t.project, t.name, e.properties_json"
}'
Python Client Usage
Using the codebase_memory_mcp package:
from codebase_memory_mcp import MCPClient
client = MCPClient()
result = client.query_graph(
project="my-repo",
query=(
"MATCH (src)-[:CROSS_HTTP_CALLS]->(tgt) "
"WHERE src.name = 'fetchUser' "
"RETURN tgt.project AS target_repo, tgt.name AS target_fn, tgt.properties_json"
)
)
print(result["results"])
Summary
- Cross-repo discovery runs as a dedicated pipeline pass in
src/pipeline/pass_cross_repo.c, scanning for matching routes and channels across all indexed projects. - Canonicalization via
cr_url_pathensures HTTP URLs match by path component only, ignoring host differences between services. - Bidirectional storage inserts
CROSS_*edges into both source and target SQLite databases, with idempotency guaranteed by unique constraints. - Rich metadata in JSON payloads includes target project names, file paths, and protocol specifics, enabling detailed traceability.
- Seamless querying through standard MCP tools (
search_graph,trace_path) treats cross-repo edges identically to local edges.
Frequently Asked Questions
What are the performance implications of cross-repo edges?
Cross-repo edges introduce negligible latency overhead because they are stored locally in each project's SQLite database. The system avoids network round-trips by opening both source and target database files during the indexing pass, then treating CROSS_* edges as standard local edges during query execution. This design ensures sub-millisecond response times for multi-repo queries comparable to intra-repo operations.
How does the system prevent duplicate cross-repo edges?
The insert_cross_edge function enforces idempotency through a unique constraint on the composite key (source_id, target_id, type). Additionally, before each indexing pass, the system executes a cleanup phase that removes all existing CROSS_* edges for the project (lines 11-18 in src/pipeline/pass_cross_repo.c). This guarantees fresh links without accumulation of stale relationships.
Can CROSS_* edges represent non-HTTP protocols?
Yes. The implementation supports multiple protocol types including CROSS_ASYNC_CALLS for asynchronous messaging, CROSS_CHANNEL for pub/sub patterns, CROSS_GRPC_CALLS for gRPC services, CROSS_GRAPHQL_CALLS for GraphQL operations, and CROSS_TRPC_CALLS for tRPC procedures. These are enumerated in src/mcp/mcp.c and handled identically to HTTP edges during the discovery pass.
Where are cross-repo edges visualized in the UI?
The HTTP server implementation in src/ui/http_server.c (lines 1318-1330) provides endpoints that fetch distinct target projects from CROSS_* edges. This enables the web interface to visualize inter-service dependencies across repository boundaries, displaying which external services consume or provide specific endpoints.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →