How Cross-Repo Search Works in Code-Review-Graph: Registry Architecture and Hybrid Ranking

Cross-repo search in code-review-graph aggregates fuzzy and exact matches across all registered repositories by loading the global registry, querying each repository's local graph database via hybrid search, and merging results while preserving per-repository relevance rankings.

The code-review-graph project enables multi-repository code analysis by maintaining a central registry of indexed codebases. Understanding how cross-repo search leverages this registry to distribute queries across isolated graph databases is essential for developers building integrated review workflows.

The Global Registry: Tracking Registered Repositories

The system persists repository metadata in ~/.code-review-graph/registry.json. The Registry class defined in code_review_graph/registry.py handles atomic read and write operations for this file. Each entry stores a filesystem path and an optional alias, allowing the search tool to locate the .code-review-graph/ subdirectory containing the SQLite graph database.

How Cross-Repo Search Executes

The cross_repo_search_func in code_review_graph/tools/registry_tools.py implements the coordination logic that fans out queries to every registered repository and aggregates the responses.

Loading and Iterating Over Repositories

The function instantiates Registry() to fetch all registered entries, short-circuiting with an empty result set if the registry contains no repositories. For each valid entry, it extracts the absolute repository path and alias, preparing to open the associated graph store.

Local Graph Database Access

For every repository entry, the function resolves the database path using get_db_path(repo_path), a helper defined in the incremental analysis module. It then initializes a GraphStore instance to open the SQLite file. These components reside in code_review_graph/graph.py and code_review_graph/incremental.py, providing the connection layer between the search tool and the persistent graph structure.

Hybrid Search Within Each Repository

With the store initialized, the function invokes hybrid_search(store, query, limit=limit_per_repo). This engine, typically implemented in code_review_graph/search/hybrid.py, combines full-text token matching with structural relevance scoring. Each result is annotated with "repo" (the alias) and "repo_path" (the absolute filesystem location) to disambiguate identical symbols across codebases.

Merging and Ranking Results

Results from individual repositories retain their local relevance order. The algorithm constructs a tuple (local_rank, repo_index, result) for every match, where repo_index reflects the registration sequence. The aggregate list sorts first by local_rank to prioritize high-confidence matches, then by repo_index to ensure deterministic ordering when scores are incomparable. The final payload includes a status flag, a human-readable summary string, the merged results array, and a repos_searched list documenting which repositories contributed matches.

Practical Implementation: Using the Search Tools

Developers can invoke cross-repo search programmatically or via the CLI entry point defined in code_review_graph/main.py.

Invoke the function directly from Python:

from code_review_graph.tools.registry_tools import cross_repo_search_func

# Query across all registered repos, limiting to 5 hits per repo

response = cross_repo_search_func(query="authentication middleware", limit=5)

print(response["summary"])

# Output: Found 8 result(s) across 2 repo(s) for 'authentication middleware'

for hit in response["results"]:
    print(f"{hit['name']} ({hit['kind']}) in {hit['repo']}: {hit['file_path']}")

Or use the command-line interface:


# Via the CLI entry point registered in code_review_graph/main.py

crg cross_repo_search "database migration"

Summary

  • Cross-repo search operates on the ~/.code-review-graph/registry.json global registry managed by code_review_graph/registry.py.
  • cross_repo_search_func iterates registered repos, opening each via GraphStore and get_db_path from code_review_graph/graph.py.
  • Each repo is queried using hybrid_search, which merges full-text and structural matching.
  • Results are merged using (local_rank, repo_index) sorting to preserve intra-repo relevance while maintaining deterministic cross-repo ordering.
  • The tool returns a structured payload including metadata about which repos were searched and how many matches were found.

Frequently Asked Questions

How does code-review-graph store repository metadata?

The system writes to ~/.code-review-graph/registry.json. The Registry class in code_review_graph/registry.py manages this JSON file, storing each repository's absolute path and optional alias.

What search algorithm does hybrid_search use?

According to the source code, hybrid_search combines full-text token matching with structural relevance scoring to rank symbols within a single repository's graph database.

How are results ranked across different repositories?

Results maintain their local relevance ranking from hybrid_search. The merger sorts by local rank first, then by repository registration index (repo_index), ensuring high-confidence matches surface first while providing deterministic tie-breaking.

Can I limit the number of results per repository?

Yes. The cross_repo_search_func accepts a limit parameter that caps the number of results returned by hybrid_search for each individual repository before the final merge step.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →