How the Multi-Hop Retrieval Benchmark Evaluates Graph Traversal in code-review-graph

The multi-hop retrieval benchmark measures graph-traversal quality through a two-step evaluation: semantic anchor discovery followed by one-hop graph traversal, tracking five metrics that combine to produce a 0-to-1 quality score.

The multi-hop retrieval benchmark in the tirth8205/code-review-graph repository provides a quantitative framework for testing how effectively an LLM-style agent can use the generated code graph to answer natural-language queries requiring relationship traversal. According to the source code, this benchmark is designed to stress-test both the semantic search capabilities and the structural accuracy of the underlying graph.

Two-Step Evaluation Architecture

The benchmark's evaluation pipeline consists of two distinct phases, each implemented in separate modules of the codebase.

Step 1: Semantic Anchor Discovery

The benchmark begins by calling hybrid_search from [code_review_graph/search.py](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/search.py) with the natural-language question. This function retrieves the top-K nodes (where K defaults to 10) and searches for a node whose qualified_name ends with the task-defined suffix specified in the YAML configuration.

Successful anchor discovery depends entirely on the quality of the embedding-based search. If the semantic search fails to surface the correct node within the top-K results, the traversal step cannot proceed meaningfully.

Step 2: One-Hop Graph Traversal

Once an anchor is identified, the benchmark invokes query_graph from [code_review_graph/tools/query.py](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/tools/query.py) with the specified traversal pattern. Supported patterns include:

  • callers_of – functions that call the anchor
  • callees_of – functions called by the anchor
  • tests_for – test cases covering the anchor

The function returns all neighbors reachable through the requested edge type, which the benchmark then compares against expected results.

Benchmark Metrics Explained

The evaluation records five specific metrics for each task defined in code_review_graph/eval/configs/*.yaml:

Metric Definition Source
anchor_found Boolean: was a matching anchor present in top-K results? hybrid_search results inspection
anchor_rank Integer position of matching node (1 = best) Result list ordering
neighbor_count Total nodes returned by query_graph Traversal output length
neighbor_recall Fraction of expected neighbors actually returned Comparison against expected_neighbor_names
score int(anchor_found) * neighbor_recall – final 0-to-1 quality measure Composite calculation

The score computation explicitly ties semantic search success to traversal accuracy: even perfect neighbor recall yields zero if the anchor wasn't found, while finding the anchor guarantees at least some non-zero score proportional to traversal correctness.

Running the Multi-Hop Retrieval Benchmark

Programmatic Execution

from pathlib import Path
from code_review_graph.eval.runner import run_benchmarks
from code_review_graph.store import Store

repo_path = Path("/path/to/your/repo")
store = Store(repo_path)

config = {
    "name": "my-repo",
    "multi_hop_tasks": [
        {
            "id": "example-1",
            "nl_query": "Where is the function that creates a user?",
            "anchor_qualified_suffix": "UserService.create_user",
            "traversal_pattern": "callers_of",
            "expected_neighbor_names": [
                "UserController.register_user",
                "UserAPI.create",
            ],
            "k": 10,
        },
    ],
    "_embedding_provider": "openai",
    "_embedding_model": "text-embedding-ada-002",
}

results = run_benchmarks.run_multi_hop(repo_path, store, config)
print(results)

Command-Line Interface

The package provides a CLI entry point defined in code_review_graph/cli.py:

code-review-graph eval --benchmark multi_hop --repo /path/to/repo

This loads YAML task definitions from code_review_graph/eval/configs/, executes the two-step process for each task, and renders a formatted table of all five metrics.

Key Implementation Files

Understanding the benchmark requires familiarity with these source files:

Why This Benchmark Design Matters

The multi-hop retrieval benchmark's two-step structure deliberately mirrors real-world usage patterns where an agent must first locate a relevant code entity conversationally, then explore its relationships. By separating anchor discovery from traversal execution, the benchmark isolates failure modes: poor scores can indicate embedding quality issues, graph construction problems, or traversal pattern limitations.

The neighbor_recall metric specifically validates that the graph edges themselves are correct—not merely that some neighbors exist, but that the expected neighbors are present. This distinguishes structural completeness from useful accuracy in the code-review-graph system.

Summary

  • The multi-hop retrieval benchmark evaluates both semantic search and graph traversal in a unified pipeline
  • Five metrics track performance from anchor discovery through neighbor recall to a composite quality score
  • Implementation spans search.py, tools/query.py, and eval/benchmarks/multi_hop_retrieval.py
  • Tasks are configured via YAML files in eval/configs/ and executable via CLI or Python API
  • The 0-to-1 score formula (int(anchor_found) * neighbor_recall) ensures no partial credit for finding wrong anchors

Frequently Asked Questions

What does "multi-hop" mean if the benchmark only performs one graph traversal?

The term describes the two-stage retrieval process: first a "hop" through embedding space to find the semantic anchor, then a structural hop through graph edges. Future iterations may extend this to multiple structural traversals, but the current implementation validates the foundational capability of combining semantic and symbolic retrieval.

How do I define custom benchmark tasks?

Create YAML files in code_review_graph/eval/configs/ specifying nl_query, anchor_qualified_suffix, traversal_pattern, and expected_neighbor_names. The benchmark loader automatically discovers and executes all valid task definitions found in this directory.

Can the benchmark run without an embedding provider configured?

No. The hybrid_search step requires an initialized embedding store with provider configuration (e.g., OpenAI, local models). The benchmark will fail during anchor discovery if embeddings are unavailable or the store is unpopulated.

What traversal patterns are supported beyond callers and callees?

As implemented in query_graph, the system supports multiple relationship types including tests_for for test coverage relationships. The exact available patterns depend on which edge types were populated during graph construction from the target repository.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →