# How opportunity_scorer.py Utilizes Position-Specific CTR Benchmarks to Score SEO Opportunities

> Discover how opportunity_scorer.py uses position-specific CTR benchmarks to score SEO opportunities. Learn how it compares actual CTR to industry averages for better insights.

- Repository: [Craig/seomachine](https://github.com/TheCraigHewitt/seomachine)
- Tags: internals
- Published: 2026-03-12

---

**The [`opportunity_scorer.py`](https://github.com/TheCraigHewitt/seomachine/blob/main/opportunity_scorer.py) module compares a keyword's actual click-through rate against position-specific industry averages stored in the `EXPECTED_CTR` dictionary, assigning higher opportunity scores when actual CTR falls significantly below the benchmark for its current SERP position.**

The `OpportunityScorer` class in the **TheCraigHewitt/seomachine** repository provides a data-driven approach to prioritizing SEO tasks by quantifying CTR underperformance relative to SERP position. By utilizing **position-specific CTR benchmarks**, the system identifies keywords where current performance lags behind typical industry expectations, flagging them as high-priority optimization candidates. This methodology allows SEO practitioners to spot "quick win" opportunities where improving title tags or meta descriptions could yield disproportionate traffic gains.

## The EXPECTED_CTR Benchmark Dictionary

At the core of the CTR evaluation logic lies the class-level constant `EXPECTED_CTR`, defined at **lines 37-59** in [`data_sources/modules/opportunity_scorer.py`](https://github.com/TheCraigHewitt/seomachine/blob/main/data_sources/modules/opportunity_scorer.py). This dictionary maps integer SERP positions (1 through 20) to their corresponding typical click-through rates based on industry aggregate data.

- **Rank 1** carries an expected CTR of approximately **31.6%**
- **Rank 12** expects roughly **1.8%**
- Positions beyond 20 default to **0.5%** (0.005) via the dictionary's `.get()` fallback mechanism

These values represent the baseline against which all actual keyword performance is judged, allowing the scorer to contextualize whether a 2% CTR is excellent (for position 15) or problematic (for position 5).

## The CTR Calculation Pipeline

When `calculate_score` is invoked, it delegates CTR analysis to the private method `_calculate_ctr_score` (**lines 74-104**). This method executes a multi-step comparison logic to quantify optimization potential.

### Position Normalization and Safe Lookup

The method first rounds the floating-point `position` value to the nearest integer (`pos_int`) to align with the benchmark keys. It then retrieves the expected CTR using:

```python
expected_ctr = self.EXPECTED_CTR.get(pos_int, 0.005)

```

This safeguards against out-of-range positions by defaulting to 0.5%. If the input data lacks a CTR field, the method automatically recomputes it from `clicks / impressions` before proceeding.

### Threshold-Based Scoring Logic

The actual CTR is compared against four percentage thresholds of the expected value: **30%**, **50%**, **70%**, and **90%**. The scoring algorithm assigns discrete values based on where the actual performance falls:

- **Below 30%** of expected → Score **100** (critical opportunity)
- **30-50%** of expected → Score **70**
- **50-70%** of expected → Score **50**
- **70-90%** of expected → Score **30**
- **Above 90%** of expected → Score **0** (performing at or above benchmark)

This inverse scoring system ensures that keywords with the largest CTR gaps receive the highest opportunity scores, directly signaling where metadata optimizations could drive immediate traffic increases.

## Weighting and Integration in the Final Score

The raw `ctr_score` generated by `_calculate_ctr_score` feeds into the composite opportunity calculation within `calculate_score` (**lines 144-152**). Here, the CTR component receives a fixed weight of **5%** (0.05) in the weighted sum alongside other SEO factors like difficulty and commercial intent.

This weighting ensures that CTR underperformance meaningfully influences the final priority ranking without overwhelming other signals. In **QUICK_WIN** opportunity analyses—where low-hanging fruit is prioritized—a low CTR score can significantly boost a keyword's overall priority despite moderate difficulty ratings.

## Practical Implementation Example

The following example demonstrates how [`opportunity_scorer.py`](https://github.com/TheCraigHewitt/seomachine/blob/main/opportunity_scorer.py) evaluates a keyword ranking at position 12 with subpar click-through performance:

```python

# Example: scoring a keyword with a low CTR at position 12

from data_sources.modules.opportunity_scorer import OpportunityScorer, OpportunityType

keyword = {
    "position": 12,            # SERP rank

    "ctr": 0.006,              # 0.6% actual CTR (well below expected ~1.8%)

    "impressions": 2000,
    "clicks": 12,
    "commercial_intent": 1.2,
}
scorer = OpportunityScorer()
result = scorer.calculate_score(
    keyword_data=keyword,
    opportunity_type=OpportunityType.QUICK_WIN,
    difficulty=35,
)

print("CTR score :", result["score_breakdown"]["ctr_score"])   # → 100 (huge opportunity)

print("Overall   :", result["final_score"])

```

Adjusting the CTR closer to the benchmark reduces the opportunity score as the performance gap narrows:

```python

# Adjust the CTR to see the effect on the CTR score

keyword["ctr"] = 0.015   # 1.5% (closer to expected 1.8%)

result = scorer.calculate_score(keyword_data=keyword,
                                opportunity_type=OpportunityType.QUICK_WIN,
                                difficulty=35)

print("CTR score :", result["score_breakdown"]["ctr_score"])   # → 70

```

Modules like [`research_quick_wins.py`](https://github.com/TheCraigHewitt/seomachine/blob/main/research_quick_wins.py) and [`research_competitor_gaps.py`](https://github.com/TheCraigHewitt/seomachine/blob/main/research_competitor_gaps.py) utilize this same scoring pipeline, demonstrating how position-specific CTR benchmarks provide consistent evaluation criteria across different SEO research workflows.

## Summary

- **`EXPECTED_CTR`** stores industry-average click-through rates for positions 1-20 at [`data_sources/modules/opportunity_scorer.py`](https://github.com/TheCraigHewitt/seomachine/blob/main/data_sources/modules/opportunity_scorer.py) lines 37-59
- **`_calculate_ctr_score`** (lines 74-104) compares actual CTR against 30%, 50%, 70%, and 90% thresholds of the position-specific benchmark
- Scores range from **100** (severe underperformance) to **0** (meeting or exceeding expectations)
- The CTR component carries a **5% weight** in the final opportunity score calculated at lines 144-152
- A **0.5% fallback CTR** protects the calculation when processing keywords ranked beyond position 20

## Frequently Asked Questions

### What data structure stores the position-specific CTR benchmarks?

The benchmarks reside in the `EXPECTED_CTR` dictionary defined at the class level in [`opportunity_scorer.py`](https://github.com/TheCraigHewitt/seomachine/blob/main/opportunity_scorer.py). This structure maps integer SERP positions (1-20) to float values representing typical click-through rates, with rank 1 set to approximately 31.6% and rank 20 set to roughly 1.0%.

### How does the scorer handle keywords with missing CTR data?

When the `ctr` key is absent from the input dictionary, `_calculate_ctr_score` automatically computes the metric by dividing `clicks` by `impressions`. If both are missing or zero, the calculation proceeds with available data, though the fallback mechanism for position lookup ensures the scoring never fails completely.

### Why does the CTR score use an inverse scale where higher numbers indicate worse performance?

The **100-to-30 scoring range** represents an "opportunity magnitude" rather than a quality grade. A score of 100 indicates the largest possible gap between expected and actual performance, signaling the highest potential ROI from optimization efforts like rewriting title tags or enhancing rich snippet eligibility.

### What happens when a keyword ranks outside the top 20 positions?

For positions beyond 20, the lookup `self.EXPECTED_CTR.get(pos_int, 0.005)` returns a default value of **0.005** (0.5%). This conservative estimate prevents IndexError exceptions while ensuring that low-ranking keywords with measurable CTR still receive appropriate scoring relative to typical page-two performance.