# MRR vs. F1: Understanding Search Quality Benchmarks for Code Retrieval

> Understand MRR vs F1 score for code retrieval search quality. Learn which benchmark suits your needs best for precise and relevant results.

- Repository: [Tirth Kanani/code-review-graph](https://github.com/tirth8205/code-review-graph)
- Tags: deep-dive
- Published: 2026-08-16

---

**Mean Reciprocal Rank (MRR) measures the rank position of the first relevant result, while F1 score balances precision and recall across all relevant items—making them complementary metrics for different search scenarios.**

Search quality benchmarks are essential for evaluating how effectively a system retrieves relevant information. The **code-review-graph** repository implements both **MRR** and **F1 score** to assess distinct aspects of code search performance. Understanding when to apply each metric helps developers optimize their retrieval systems for specific user goals.

## What MRR Measures: Ranking Quality for Single-Answer Scenarios

**Mean Reciprocal Rank (MRR)** evaluates how quickly a search system surfaces the first correct result. It is calculated as the average of reciprocal ranks across multiple queries.

```python

# Example: computing MRR for a list of query-to-result rankings

from code_review_graph.eval.scorer import mean_reciprocal_rank

# ranks[i] = position (1-based) of the first relevant result for query i

ranks = [1, 3, 2, 5]      # sample data

mrr = mean_reciprocal_rank(ranks)
print(f"Mean Reciprocal Rank: {mrr:.3f}")   # → 0.458

```

**Key characteristics of MRR:**

- **Range:** 0 to 1, where 1 indicates every query's first result is relevant
- **Sensitivity:** Extremely sensitive to top-ranked position—a result at rank 2 contributes only half as much as rank 1
- **Best for:** Single-answer lookup scenarios where users want the correct hit immediately

In [`code_review_graph/eval/benchmarks/search_quality.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/eval/benchmarks/search_quality.py), MRR serves as the primary metric for assessing graph-based code element retrieval. The reporter module in [`code_review_graph/eval/reporter.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/eval/reporter.py) outputs a dedicated "Search MRR" column in benchmark reports.

## What F1 Measures: Balanced Retrieval for Set-Based Results

The **F1 score** combines **precision** (relevant items retrieved ÷ total items retrieved) and **recall** (relevant items retrieved ÷ total relevant items) into a single harmonic mean.

```python

# Example: computing Precision, Recall, and F1 for a set-retrieval task

from code_review_graph.eval.scorer import precision_recall_f1

ground_truth = {"a.py:func1", "b.js:varX", "c.rb:ClassY"}
retrieved    = {"a.py:func1", "b.js:varX", "d.go:FuncZ"}

prec, rec, f1 = precision_recall_f1(retrieved, ground_truth)
print(f"Precision: {prec:.2f}, Recall: {rec:.2f}, F1: {f1:.2f}")

# → Precision: 0.67, Recall: 0.67, F1: 0.67

```

**Key characteristics of F1:**

- **Range:** 0 to 1, where 1 indicates perfect precision and recall
- **Balance:** Tolerates trade-offs between false positives and false negatives
- **Best for:** Set-based retrieval where multiple correct answers exist

The generic scorer in [`code_review_graph/eval/scorer.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/eval/scorer.py) implements F1 calculation for tasks like flow-recall and impact-derived edge detection—scenarios where completeness matters as much as accuracy.

## When to Use MRR vs. F1 in Code Search

| Use Case | Recommended Metric | Rationale |
|----------|-------------------|-----------|
| Finding a specific function or class definition | **MRR** | Users need the correct result at position 1 |
| Retrieving all impacted code from a change | **F1** | Multiple relevant files must be found completely |
| FAQ or documentation lookup | **MRR** | Single correct answer expected |
| Code review suggestion generation | **F1** | Suggestions are ranked lists requiring coverage |

According to the code-review-graph source code, these metrics are not mutually exclusive. The evaluation suite reports both to give developers a complete picture: MRR reveals ranking effectiveness, while F1 exposes retrieval completeness.

## Implementation Details in code-review-graph

Three core files handle these search quality benchmarks:

- **[`code_review_graph/eval/scorer.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/eval/scorer.py)** — Implements `mean_reciprocal_rank()` and `precision_recall_f1()` functions
- **[`code_review_graph/eval/benchmarks/search_quality.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/eval/benchmarks/search_quality.py)** — Defines MRR-driven search quality evaluation
- **[`code_review_graph/eval/reporter.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/eval/reporter.py)** — Formats output tables with "Search MRR" as a primary column

The repository also includes visualization support in [`diagrams/generate_diagrams.py`](https://github.com/tirth8205/code-review-graph/blob/main/diagrams/generate_diagrams.py) for plotting recall-versus-F1 trade-offs.

## Summary

- **MRR** optimizes for speed-to-first-relevant-result; ideal when users need one correct answer immediately
- **F1** optimizes for balanced retrieval quality; essential when multiple relevant items must be found
- The code-review-graph repository implements both in [`scorer.py`](https://github.com/tirth8205/code-review-graph/blob/main/scorer.py), applying MRR for ranking evaluation and F1 for set-based tasks
- Select your metric based on user intent: lookup scenarios favor MRR, comprehensive retrieval favors F1

## Frequently Asked Questions

### Can MRR and F1 be used together?

Yes. Many retrieval systems report both metrics to capture different success dimensions. MRR indicates whether users see relevant items quickly, while F1 shows whether the system finds all relevant items without excessive noise.

### Why does MRR drop so sharply when the first relevant result moves down one position?

MRR uses reciprocal rank (1/rank), so rank 1 contributes 1.0, rank 2 contributes 0.5, and rank 3 contributes only 0.33. This steep decline reflects the real-world behavior of search users, who rarely examine results beyond the top few positions.

### Is F1 always better than using precision or recall alone?

F1 provides a balanced view but hides trade-offs. A system with 100% precision and 0% recall yields F1 = 0, as does 0% precision and 100% recall. When the costs of false positives and false negatives differ substantially, inspect precision and recall separately rather than relying solely on F1.

### How does code-review-graph handle cases with no relevant results?

The scorer implementations in [`code_review_graph/eval/scorer.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/eval/scorer.py) typically assign zero contribution for queries without relevant results in MRR calculations, and treat empty ground truth sets as edge cases requiring explicit handling in F1 computation.