How Cross-Repo Intelligence Generates CROSS_* Edges in Codebase Memory
Cross-repo intelligence in Codebase Memory works by scanning SQLite stores across multiple repositories, matching call edges to route definitions, and emitting bidirectional CROSS_* edges that link callers and callees across project boundaries.
Codebase Memory implements cross-repo intelligence through a dedicated pipeline that transforms intra-project call edges into inter-project links. This system enables the graph database to reason about interactions spanning multiple repositories by generating special edge types prefixed with CROSS_. The entire mechanism lives in the cross-repo pipeline implementation within the DeusData/codebase-memory-mcp repository.
The Cross-Repo Pipeline Architecture
The cross-repo matching system centers on a single public API that orchestrates store scanning, edge matching, and bidirectional link creation.
Core API Entry Point
All cross-repo operations begin with the cbm_cross_repo_match() function defined in src/pipeline/pass_cross_repo.c (lines 751-770). This function accepts a source project name, an array of target projects, and returns a cbm_cross_repo_result_t structure containing statistics for each edge type created.
cbm_cross_repo_result_t cbm_cross_repo_match(const char *project,
const char **target_projects,
int target_count);
Project Discovery and Store Management
The pipeline first opens the source project's SQLite store and clears any stale CROSS_* edges to ensure each run works from a clean slate (lines 561-569). When the caller passes "*" as a target, the helper collect_all_projects() scans the cache directory for every *.db file while ignoring internal files such as *_cross_repo.db or SQLite-WAL files (lines 696-718).
The function then iterates over each target project (skipping the source itself) and opens its store read-write (lines 783-795).
How CROSS_* Edges Are Generated
The system matches four distinct categories of communication patterns between repositories, each generating specific CROSS_* edge types.
HTTP Route Matching (CROSS_HTTP_CALLS)
The match_http_routes() function reads HTTP_CALLS edges from the source store and canonicalizes the caller URL using cr_url_path() and cbm_route_canon_path(). It builds a qualified route name in the format __route__<METHOD>__<path> and attempts an exact lookup; if that fails, it falls back to template matching via find_route_handler_fuzzy() (lines 889-914).
When a handler node is found, emit_cross_route_bidirectional() creates a CROSS_HTTP_CALLS edge with properties including target_project, url_path, and method. The system inserts this edge twice—once forward (source → target) and once reverse (target → source)—ensuring visibility from either repository (lines 443-458).
Async Topic Matching (CROSS_ASYNC_CALLS)
The match_async_routes() function handles ASYNC_CALLS edges by building a qualified name __route__<broker>__<url_path>. It inserts a CROSS_ASYNC_CALLS edge forward only (lines 508-514), as the reverse direction is covered by the symmetric HTTP pass.
Channel Communication (CROSS_CHANNEL)
For message channel patterns, match_channels() enumerates EMITS edges on Channel nodes. The helper try_match_channel_listener() searches the target store for a LISTENS_ON edge with the same channel name. When both ends exist, the system adds a CROSS_CHANNEL edge forward (emitter → channel) and reverse (listener → target channel) with properties including target_project, channel_name, and transport (lines 625-647, 486-518).
Typed RPC Protocols (CROSS_GRPC_CALLS, CROSS_GRAPHQL_CALLS, CROSS_TRPC_CALLS)
The generic match_typed_routes() function handles gRPC, GraphQL, and tRPC protocols. It extracts service- and method-specific JSON fields (service, method, etc.), looks up the target route node, and emits the corresponding CROSS_GRPC_CALLS, CROSS_GRAPHQL_CALLS, or CROSS_TRPC_CALLS edge using the same bidirectional helper (lines 822-842).
Edge Insertion and Idempotency
Each matcher builds the CROSS_* edge properties JSON and delegates to insert_cross_edge(), a thin wrapper around cbm_store_insert_edge(). This routine upserts on the unique (source_id, target_id, type) key, guaranteeing idempotency (lines 121-135). Because insertion is idempotent, repeated runs after code changes safely update the graph without duplication.
After processing each target, the system aggregates counts of created edges (result.http_edges, result.async_edges, etc.) and records elapsed time in milliseconds (lines 1276-1284). Finally, both source and target stores are closed and temporary allocations are freed (lines 1800-1812).
Practical Implementation Example
To run cross-repo matching for a project against all other repositories:
/* Example: run cross-repo matching for project "frontend" against all other projects */
const char *targets[] = { "*" };
cbm_cross_repo_result_t stats = cbm_cross_repo_match("frontend", targets, 1);
/* stats now contains the number of each CROSS_* edge type that was created */
printf("Created %d CROSS_HTTP_CALLS edges, %d CROSS_CHANNEL edges\n",
stats.http_edges, stats.channel_edges);
For integration with higher-level commands, such as the codebase-memory-mcp binary:
int main(int argc, char **argv) {
if (argc != 2) {
fprintf(stderr, "Usage: %s <project>\n", argv[0]);
return 1;
}
const char *proj = argv[1];
const char *all[] = { "*" };
cbm_cross_repo_result_t r = cbm_cross_repo_match(proj, all, 1);
cbm_log_info("cross_repo.done", "project", proj, "total_cross_edges",
cbm_itoa(r.http_edges + r.async_edges + r.channel_edges +
r.grpc_edges + r.graphql_edges + r.trpc_edges));
return 0;
}
Summary
- Cross-repo intelligence generates
CROSS_*edges by matching concrete call edges in one project to route definitions in another. - The pipeline resides in
src/pipeline/pass_cross_repo.cand exposes thecbm_cross_repo_match()API. - Six edge types are supported:
CROSS_HTTP_CALLS,CROSS_ASYNC_CALLS,CROSS_CHANNEL,CROSS_GRPC_CALLS,CROSS_GRAPHQL_CALLS, andCROSS_TRPC_CALLS. - Bidirectional insertion ensures links are visible from both source and target projects.
- Idempotent upserts prevent edge duplication across multiple runs.
- Project discovery supports wildcard (
"*") scanning of the cache directory for all*.dbfiles.
Frequently Asked Questions
What is the purpose of CROSS_* edges in Codebase Memory?
CROSS_* edges represent inter-project dependencies that span repository boundaries. They allow the graph to model HTTP calls, async messages, channel communications, and RPC invocations where the caller and callee reside in different codebases, enabling comprehensive architectural analysis across microservices.
How does the system prevent duplicate CROSS_* edges when running repeatedly?
The insert_cross_edge() function in src/pipeline/pass_cross_repo.c performs an upsert operation using the unique composite key (source_id, target_id, type). This guarantees idempotency—running the cross-repo matcher multiple times updates existing edges rather than creating duplicates.
Can cross-repo matching target specific projects rather than all repositories?
Yes. While passing "*" to cbm_cross_repo_match() scans all projects in the cache directory, you can provide specific project names in the target_projects array. The function skips the source project automatically and only matches against the explicitly listed targets.
Where are CROSS_* edges queried after creation?
The HTTP server in src/ui/http_server.c provides endpoints that query CROSS_% edges using SQL patterns like SELECT … WHERE type LIKE 'CROSS_%'. Additionally, src/pipeline/pipeline.c adds cross-repo link summaries to the architecture JSON output when any CROSS_* edges exist in the graph.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →