# How the Multi-Hop Retrieval Benchmark Evaluates Graph Traversal in code-review-graph

> Evaluate graph traversal quality in code-review-graph using the multi-hop retrieval benchmark. Discover semantic anchors and track five metrics for a 0-to-1 quality score.

- Repository: [Tirth Kanani/code-review-graph](https://github.com/tirth8205/code-review-graph)
- Tags: benchmark
- Published: 2026-08-15

---

**The multi-hop retrieval benchmark measures graph-traversal quality through a two-step evaluation: semantic anchor discovery followed by one-hop graph traversal, tracking five metrics that combine to produce a 0-to-1 quality score.**

The **multi-hop retrieval benchmark** in the `tirth8205/code-review-graph` repository provides a quantitative framework for testing how effectively an LLM-style agent can use the generated code graph to answer natural-language queries requiring relationship traversal. According to the source code, this benchmark is designed to stress-test both the semantic search capabilities and the structural accuracy of the underlying graph.

## Two-Step Evaluation Architecture

The benchmark's evaluation pipeline consists of two distinct phases, each implemented in separate modules of the codebase.

### Step 1: Semantic Anchor Discovery

The benchmark begins by calling `hybrid_search` from [[`code_review_graph/search.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/search.py)](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/search.py) with the natural-language question. This function retrieves the top-K nodes (where **K defaults to 10**) and searches for a node whose `qualified_name` ends with the task-defined suffix specified in the YAML configuration.

Successful anchor discovery depends entirely on the quality of the embedding-based search. If the semantic search fails to surface the correct node within the top-K results, the traversal step cannot proceed meaningfully.

### Step 2: One-Hop Graph Traversal

Once an anchor is identified, the benchmark invokes `query_graph` from [[`code_review_graph/tools/query.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/tools/query.py)](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/tools/query.py) with the specified traversal pattern. Supported patterns include:

- `callers_of` – functions that call the anchor
- `callees_of` – functions called by the anchor
- `tests_for` – test cases covering the anchor

The function returns all neighbors reachable through the requested edge type, which the benchmark then compares against expected results.

## Benchmark Metrics Explained

The evaluation records five specific metrics for each task defined in `code_review_graph/eval/configs/*.yaml`:

| Metric | Definition | Source |
|--------|-----------|--------|
| **anchor_found** | Boolean: was a matching anchor present in top-K results? | `hybrid_search` results inspection |
| **anchor_rank** | Integer position of matching node (1 = best) | Result list ordering |
| **neighbor_count** | Total nodes returned by `query_graph` | Traversal output length |
| **neighbor_recall** | Fraction of expected neighbors actually returned | Comparison against `expected_neighbor_names` |
| **score** | `int(anchor_found) * neighbor_recall` – final 0-to-1 quality measure | Composite calculation |

The **score** computation explicitly ties semantic search success to traversal accuracy: even perfect neighbor recall yields zero if the anchor wasn't found, while finding the anchor guarantees at least some non-zero score proportional to traversal correctness.

## Running the Multi-Hop Retrieval Benchmark

### Programmatic Execution

```python
from pathlib import Path
from code_review_graph.eval.runner import run_benchmarks
from code_review_graph.store import Store

repo_path = Path("/path/to/your/repo")
store = Store(repo_path)

config = {
    "name": "my-repo",
    "multi_hop_tasks": [
        {
            "id": "example-1",
            "nl_query": "Where is the function that creates a user?",
            "anchor_qualified_suffix": "UserService.create_user",
            "traversal_pattern": "callers_of",
            "expected_neighbor_names": [
                "UserController.register_user",
                "UserAPI.create",
            ],
            "k": 10,
        },
    ],
    "_embedding_provider": "openai",
    "_embedding_model": "text-embedding-ada-002",
}

results = run_benchmarks.run_multi_hop(repo_path, store, config)
print(results)

```

### Command-Line Interface

The package provides a CLI entry point defined in [`code_review_graph/cli.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/cli.py):

```bash
code-review-graph eval --benchmark multi_hop --repo /path/to/repo

```

This loads YAML task definitions from `code_review_graph/eval/configs/`, executes the two-step process for each task, and renders a formatted table of all five metrics.

## Key Implementation Files

Understanding the benchmark requires familiarity with these source files:

- **[`code_review_graph/eval/benchmarks/multi_hop_retrieval.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/eval/benchmarks/multi_hop_retrieval.py)** – Core benchmark logic implementing anchor search, graph traversal, and metric calculation
- **[`code_review_graph/search.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/search.py)** – `hybrid_search` implementation for semantic node retrieval
- **[`code_review_graph/tools/query.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/tools/query.py)** – `query_graph` implementation for pattern-based traversal
- **`code_review_graph/eval/configs/*.yaml`** – Task definitions specifying queries, expected anchors, and expected neighbors
- **[`code_review_graph/cli.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/cli.py)** – CLI entry point for benchmark invocation

## Why This Benchmark Design Matters

The multi-hop retrieval benchmark's two-step structure deliberately mirrors real-world usage patterns where an agent must first locate a relevant code entity conversationally, then explore its relationships. By separating **anchor discovery** from **traversal execution**, the benchmark isolates failure modes: poor scores can indicate embedding quality issues, graph construction problems, or traversal pattern limitations.

The `neighbor_recall` metric specifically validates that the graph edges themselves are correct—not merely that some neighbors exist, but that the *expected* neighbors are present. This distinguishes structural completeness from useful accuracy in the `code-review-graph` system.

## Summary

- The multi-hop retrieval benchmark evaluates **both semantic search and graph traversal** in a unified pipeline
- **Five metrics** track performance from anchor discovery through neighbor recall to a composite quality score
- Implementation spans [`search.py`](https://github.com/tirth8205/code-review-graph/blob/main/search.py), [`tools/query.py`](https://github.com/tirth8205/code-review-graph/blob/main/tools/query.py), and [`eval/benchmarks/multi_hop_retrieval.py`](https://github.com/tirth8205/code-review-graph/blob/main/eval/benchmarks/multi_hop_retrieval.py)
- Tasks are configured via YAML files in `eval/configs/` and executable via CLI or Python API
- The 0-to-1 **score** formula (`int(anchor_found) * neighbor_recall`) ensures no partial credit for finding wrong anchors

## Frequently Asked Questions

### What does "multi-hop" mean if the benchmark only performs one graph traversal?

The term describes the **two-stage retrieval process**: first a "hop" through embedding space to find the semantic anchor, then a structural hop through graph edges. Future iterations may extend this to multiple structural traversals, but the current implementation validates the foundational capability of combining semantic and symbolic retrieval.

### How do I define custom benchmark tasks?

Create YAML files in `code_review_graph/eval/configs/` specifying `nl_query`, `anchor_qualified_suffix`, `traversal_pattern`, and `expected_neighbor_names`. The benchmark loader automatically discovers and executes all valid task definitions found in this directory.

### Can the benchmark run without an embedding provider configured?

No. The `hybrid_search` step requires an initialized embedding store with provider configuration (e.g., OpenAI, local models). The benchmark will fail during anchor discovery if embeddings are unavailable or the store is unpopulated.

### What traversal patterns are supported beyond callers and callees?

As implemented in `query_graph`, the system supports multiple relationship types including `tests_for` for test coverage relationships. The exact available patterns depend on which edge types were populated during graph construction from the target repository.