How the Multi-Hop Retrieval Benchmark Evaluates Graph Traversal in code-review-graph
The multi-hop retrieval benchmark measures graph-traversal quality through a two-step evaluation: semantic anchor discovery followed by one-hop graph traversal, tracking five metrics that combine to produce a 0-to-1 quality score.
The multi-hop retrieval benchmark in the tirth8205/code-review-graph repository provides a quantitative framework for testing how effectively an LLM-style agent can use the generated code graph to answer natural-language queries requiring relationship traversal. According to the source code, this benchmark is designed to stress-test both the semantic search capabilities and the structural accuracy of the underlying graph.
Two-Step Evaluation Architecture
The benchmark's evaluation pipeline consists of two distinct phases, each implemented in separate modules of the codebase.
Step 1: Semantic Anchor Discovery
The benchmark begins by calling hybrid_search from [code_review_graph/search.py](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/search.py) with the natural-language question. This function retrieves the top-K nodes (where K defaults to 10) and searches for a node whose qualified_name ends with the task-defined suffix specified in the YAML configuration.
Successful anchor discovery depends entirely on the quality of the embedding-based search. If the semantic search fails to surface the correct node within the top-K results, the traversal step cannot proceed meaningfully.
Step 2: One-Hop Graph Traversal
Once an anchor is identified, the benchmark invokes query_graph from [code_review_graph/tools/query.py](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/tools/query.py) with the specified traversal pattern. Supported patterns include:
callers_of– functions that call the anchorcallees_of– functions called by the anchortests_for– test cases covering the anchor
The function returns all neighbors reachable through the requested edge type, which the benchmark then compares against expected results.
Benchmark Metrics Explained
The evaluation records five specific metrics for each task defined in code_review_graph/eval/configs/*.yaml:
| Metric | Definition | Source |
|---|---|---|
| anchor_found | Boolean: was a matching anchor present in top-K results? | hybrid_search results inspection |
| anchor_rank | Integer position of matching node (1 = best) | Result list ordering |
| neighbor_count | Total nodes returned by query_graph |
Traversal output length |
| neighbor_recall | Fraction of expected neighbors actually returned | Comparison against expected_neighbor_names |
| score | int(anchor_found) * neighbor_recall – final 0-to-1 quality measure |
Composite calculation |
The score computation explicitly ties semantic search success to traversal accuracy: even perfect neighbor recall yields zero if the anchor wasn't found, while finding the anchor guarantees at least some non-zero score proportional to traversal correctness.
Running the Multi-Hop Retrieval Benchmark
Programmatic Execution
from pathlib import Path
from code_review_graph.eval.runner import run_benchmarks
from code_review_graph.store import Store
repo_path = Path("/path/to/your/repo")
store = Store(repo_path)
config = {
"name": "my-repo",
"multi_hop_tasks": [
{
"id": "example-1",
"nl_query": "Where is the function that creates a user?",
"anchor_qualified_suffix": "UserService.create_user",
"traversal_pattern": "callers_of",
"expected_neighbor_names": [
"UserController.register_user",
"UserAPI.create",
],
"k": 10,
},
],
"_embedding_provider": "openai",
"_embedding_model": "text-embedding-ada-002",
}
results = run_benchmarks.run_multi_hop(repo_path, store, config)
print(results)
Command-Line Interface
The package provides a CLI entry point defined in code_review_graph/cli.py:
code-review-graph eval --benchmark multi_hop --repo /path/to/repo
This loads YAML task definitions from code_review_graph/eval/configs/, executes the two-step process for each task, and renders a formatted table of all five metrics.
Key Implementation Files
Understanding the benchmark requires familiarity with these source files:
code_review_graph/eval/benchmarks/multi_hop_retrieval.py– Core benchmark logic implementing anchor search, graph traversal, and metric calculationcode_review_graph/search.py–hybrid_searchimplementation for semantic node retrievalcode_review_graph/tools/query.py–query_graphimplementation for pattern-based traversalcode_review_graph/eval/configs/*.yaml– Task definitions specifying queries, expected anchors, and expected neighborscode_review_graph/cli.py– CLI entry point for benchmark invocation
Why This Benchmark Design Matters
The multi-hop retrieval benchmark's two-step structure deliberately mirrors real-world usage patterns where an agent must first locate a relevant code entity conversationally, then explore its relationships. By separating anchor discovery from traversal execution, the benchmark isolates failure modes: poor scores can indicate embedding quality issues, graph construction problems, or traversal pattern limitations.
The neighbor_recall metric specifically validates that the graph edges themselves are correct—not merely that some neighbors exist, but that the expected neighbors are present. This distinguishes structural completeness from useful accuracy in the code-review-graph system.
Summary
- The multi-hop retrieval benchmark evaluates both semantic search and graph traversal in a unified pipeline
- Five metrics track performance from anchor discovery through neighbor recall to a composite quality score
- Implementation spans
search.py,tools/query.py, andeval/benchmarks/multi_hop_retrieval.py - Tasks are configured via YAML files in
eval/configs/and executable via CLI or Python API - The 0-to-1 score formula (
int(anchor_found) * neighbor_recall) ensures no partial credit for finding wrong anchors
Frequently Asked Questions
What does "multi-hop" mean if the benchmark only performs one graph traversal?
The term describes the two-stage retrieval process: first a "hop" through embedding space to find the semantic anchor, then a structural hop through graph edges. Future iterations may extend this to multiple structural traversals, but the current implementation validates the foundational capability of combining semantic and symbolic retrieval.
How do I define custom benchmark tasks?
Create YAML files in code_review_graph/eval/configs/ specifying nl_query, anchor_qualified_suffix, traversal_pattern, and expected_neighbor_names. The benchmark loader automatically discovers and executes all valid task definitions found in this directory.
Can the benchmark run without an embedding provider configured?
No. The hybrid_search step requires an initialized embedding store with provider configuration (e.g., OpenAI, local models). The benchmark will fail during anchor discovery if embeddings are unavailable or the store is unpopulated.
What traversal patterns are supported beyond callers and callees?
As implemented in query_graph, the system supports multiple relationship types including tests_for for test coverage relationships. The exact available patterns depend on which edge types were populated during graph construction from the target repository.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →