# How Cross-Repo Search Works in Code-Review-Graph: Registry Architecture and Hybrid Ranking

> Discover how cross-repo search in code-review-graph aggregates fuzzy and exact matches across all registered repositories using registry architecture and hybrid ranking.

- Repository: [Tirth Kanani/code-review-graph](https://github.com/tirth8205/code-review-graph)
- Tags: architecture
- Published: 2026-08-14

---

**Cross-repo search in code-review-graph aggregates fuzzy and exact matches across all registered repositories by loading the global registry, querying each repository's local graph database via hybrid search, and merging results while preserving per-repository relevance rankings.**

The code-review-graph project enables multi-repository code analysis by maintaining a central registry of indexed codebases. Understanding how cross-repo search leverages this registry to distribute queries across isolated graph databases is essential for developers building integrated review workflows.

## The Global Registry: Tracking Registered Repositories

The system persists repository metadata in `~/.code-review-graph/registry.json`. The `Registry` class defined in [`code_review_graph/registry.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/registry.py) handles atomic read and write operations for this file. Each entry stores a filesystem path and an optional alias, allowing the search tool to locate the `.code-review-graph/` subdirectory containing the SQLite graph database.

## How Cross-Repo Search Executes

The `cross_repo_search_func` in [`code_review_graph/tools/registry_tools.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/tools/registry_tools.py) implements the coordination logic that fans out queries to every registered repository and aggregates the responses.

### Loading and Iterating Over Repositories

The function instantiates `Registry()` to fetch all registered entries, short-circuiting with an empty result set if the registry contains no repositories. For each valid entry, it extracts the absolute repository path and alias, preparing to open the associated graph store.

### Local Graph Database Access

For every repository entry, the function resolves the database path using `get_db_path(repo_path)`, a helper defined in the incremental analysis module. It then initializes a `GraphStore` instance to open the SQLite file. These components reside in [`code_review_graph/graph.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/graph.py) and [`code_review_graph/incremental.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/incremental.py), providing the connection layer between the search tool and the persistent graph structure.

### Hybrid Search Within Each Repository

With the store initialized, the function invokes `hybrid_search(store, query, limit=limit_per_repo)`. This engine, typically implemented in [`code_review_graph/search/hybrid.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/search/hybrid.py), combines full-text token matching with structural relevance scoring. Each result is annotated with `"repo"` (the alias) and `"repo_path"` (the absolute filesystem location) to disambiguate identical symbols across codebases.

### Merging and Ranking Results

Results from individual repositories retain their local relevance order. The algorithm constructs a tuple `(local_rank, repo_index, result)` for every match, where `repo_index` reflects the registration sequence. The aggregate list sorts first by `local_rank` to prioritize high-confidence matches, then by `repo_index` to ensure deterministic ordering when scores are incomparable. The final payload includes a status flag, a human-readable summary string, the merged `results` array, and a `repos_searched` list documenting which repositories contributed matches.

## Practical Implementation: Using the Search Tools

Developers can invoke cross-repo search programmatically or via the CLI entry point defined in [`code_review_graph/main.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/main.py).

Invoke the function directly from Python:

```python
from code_review_graph.tools.registry_tools import cross_repo_search_func

# Query across all registered repos, limiting to 5 hits per repo

response = cross_repo_search_func(query="authentication middleware", limit=5)

print(response["summary"])

# Output: Found 8 result(s) across 2 repo(s) for 'authentication middleware'

for hit in response["results"]:
    print(f"{hit['name']} ({hit['kind']}) in {hit['repo']}: {hit['file_path']}")

```

Or use the command-line interface:

```bash

# Via the CLI entry point registered in code_review_graph/main.py

crg cross_repo_search "database migration"

```

## Summary

- Cross-repo search operates on the `~/.code-review-graph/registry.json` global registry managed by [`code_review_graph/registry.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/registry.py).
- `cross_repo_search_func` iterates registered repos, opening each via `GraphStore` and `get_db_path` from [`code_review_graph/graph.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/graph.py).
- Each repo is queried using `hybrid_search`, which merges full-text and structural matching.
- Results are merged using `(local_rank, repo_index)` sorting to preserve intra-repo relevance while maintaining deterministic cross-repo ordering.
- The tool returns a structured payload including metadata about which repos were searched and how many matches were found.

## Frequently Asked Questions

### How does code-review-graph store repository metadata?

The system writes to `~/.code-review-graph/registry.json`. The `Registry` class in [`code_review_graph/registry.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/registry.py) manages this JSON file, storing each repository's absolute path and optional alias.

### What search algorithm does hybrid_search use?

According to the source code, `hybrid_search` combines full-text token matching with structural relevance scoring to rank symbols within a single repository's graph database.

### How are results ranked across different repositories?

Results maintain their local relevance ranking from `hybrid_search`. The merger sorts by local rank first, then by repository registration index (`repo_index`), ensuring high-confidence matches surface first while providing deterministic tie-breaking.

### Can I limit the number of results per repository?

Yes. The `cross_repo_search_func` accepts a `limit` parameter that caps the number of results returned by `hybrid_search` for each individual repository before the final merge step.