MRR vs. F1: Understanding Search Quality Benchmarks for Code Retrieval

Mean Reciprocal Rank (MRR) measures the rank position of the first relevant result, while F1 score balances precision and recall across all relevant items—making them complementary metrics for different search scenarios.

Search quality benchmarks are essential for evaluating how effectively a system retrieves relevant information. The code-review-graph repository implements both MRR and F1 score to assess distinct aspects of code search performance. Understanding when to apply each metric helps developers optimize their retrieval systems for specific user goals.

What MRR Measures: Ranking Quality for Single-Answer Scenarios

Mean Reciprocal Rank (MRR) evaluates how quickly a search system surfaces the first correct result. It is calculated as the average of reciprocal ranks across multiple queries.


# Example: computing MRR for a list of query-to-result rankings

from code_review_graph.eval.scorer import mean_reciprocal_rank

# ranks[i] = position (1-based) of the first relevant result for query i

ranks = [1, 3, 2, 5]      # sample data

mrr = mean_reciprocal_rank(ranks)
print(f"Mean Reciprocal Rank: {mrr:.3f}")   # → 0.458

Key characteristics of MRR:

  • Range: 0 to 1, where 1 indicates every query's first result is relevant
  • Sensitivity: Extremely sensitive to top-ranked position—a result at rank 2 contributes only half as much as rank 1
  • Best for: Single-answer lookup scenarios where users want the correct hit immediately

In code_review_graph/eval/benchmarks/search_quality.py, MRR serves as the primary metric for assessing graph-based code element retrieval. The reporter module in code_review_graph/eval/reporter.py outputs a dedicated "Search MRR" column in benchmark reports.

What F1 Measures: Balanced Retrieval for Set-Based Results

The F1 score combines precision (relevant items retrieved ÷ total items retrieved) and recall (relevant items retrieved ÷ total relevant items) into a single harmonic mean.


# Example: computing Precision, Recall, and F1 for a set-retrieval task

from code_review_graph.eval.scorer import precision_recall_f1

ground_truth = {"a.py:func1", "b.js:varX", "c.rb:ClassY"}
retrieved    = {"a.py:func1", "b.js:varX", "d.go:FuncZ"}

prec, rec, f1 = precision_recall_f1(retrieved, ground_truth)
print(f"Precision: {prec:.2f}, Recall: {rec:.2f}, F1: {f1:.2f}")

# → Precision: 0.67, Recall: 0.67, F1: 0.67

Key characteristics of F1:

  • Range: 0 to 1, where 1 indicates perfect precision and recall
  • Balance: Tolerates trade-offs between false positives and false negatives
  • Best for: Set-based retrieval where multiple correct answers exist

The generic scorer in code_review_graph/eval/scorer.py implements F1 calculation for tasks like flow-recall and impact-derived edge detection—scenarios where completeness matters as much as accuracy.

Use Case Recommended Metric Rationale
Finding a specific function or class definition MRR Users need the correct result at position 1
Retrieving all impacted code from a change F1 Multiple relevant files must be found completely
FAQ or documentation lookup MRR Single correct answer expected
Code review suggestion generation F1 Suggestions are ranked lists requiring coverage

According to the code-review-graph source code, these metrics are not mutually exclusive. The evaluation suite reports both to give developers a complete picture: MRR reveals ranking effectiveness, while F1 exposes retrieval completeness.

Implementation Details in code-review-graph

Three core files handle these search quality benchmarks:

The repository also includes visualization support in diagrams/generate_diagrams.py for plotting recall-versus-F1 trade-offs.

Summary

  • MRR optimizes for speed-to-first-relevant-result; ideal when users need one correct answer immediately
  • F1 optimizes for balanced retrieval quality; essential when multiple relevant items must be found
  • The code-review-graph repository implements both in scorer.py, applying MRR for ranking evaluation and F1 for set-based tasks
  • Select your metric based on user intent: lookup scenarios favor MRR, comprehensive retrieval favors F1

Frequently Asked Questions

Can MRR and F1 be used together?

Yes. Many retrieval systems report both metrics to capture different success dimensions. MRR indicates whether users see relevant items quickly, while F1 shows whether the system finds all relevant items without excessive noise.

Why does MRR drop so sharply when the first relevant result moves down one position?

MRR uses reciprocal rank (1/rank), so rank 1 contributes 1.0, rank 2 contributes 0.5, and rank 3 contributes only 0.33. This steep decline reflects the real-world behavior of search users, who rarely examine results beyond the top few positions.

Is F1 always better than using precision or recall alone?

F1 provides a balanced view but hides trade-offs. A system with 100% precision and 0% recall yields F1 = 0, as does 0% precision and 100% recall. When the costs of false positives and false negatives differ substantially, inspect precision and recall separately rather than relying solely on F1.

How does code-review-graph handle cases with no relevant results?

The scorer implementations in code_review_graph/eval/scorer.py typically assign zero contribution for queries without relevant results in MRR calculations, and treat empty ground truth sets as edge cases requiring explicit handling in F1 computation.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →