How opportunity_scorer.py Utilizes Position-Specific CTR Benchmarks to Score SEO Opportunities

The opportunity_scorer.py module compares a keyword's actual click-through rate against position-specific industry averages stored in the EXPECTED_CTR dictionary, assigning higher opportunity scores when actual CTR falls significantly below the benchmark for its current SERP position.

The OpportunityScorer class in the TheCraigHewitt/seomachine repository provides a data-driven approach to prioritizing SEO tasks by quantifying CTR underperformance relative to SERP position. By utilizing position-specific CTR benchmarks, the system identifies keywords where current performance lags behind typical industry expectations, flagging them as high-priority optimization candidates. This methodology allows SEO practitioners to spot "quick win" opportunities where improving title tags or meta descriptions could yield disproportionate traffic gains.

The EXPECTED_CTR Benchmark Dictionary

At the core of the CTR evaluation logic lies the class-level constant EXPECTED_CTR, defined at lines 37-59 in data_sources/modules/opportunity_scorer.py. This dictionary maps integer SERP positions (1 through 20) to their corresponding typical click-through rates based on industry aggregate data.

  • Rank 1 carries an expected CTR of approximately 31.6%
  • Rank 12 expects roughly 1.8%
  • Positions beyond 20 default to 0.5% (0.005) via the dictionary's .get() fallback mechanism

These values represent the baseline against which all actual keyword performance is judged, allowing the scorer to contextualize whether a 2% CTR is excellent (for position 15) or problematic (for position 5).

The CTR Calculation Pipeline

When calculate_score is invoked, it delegates CTR analysis to the private method _calculate_ctr_score (lines 74-104). This method executes a multi-step comparison logic to quantify optimization potential.

Position Normalization and Safe Lookup

The method first rounds the floating-point position value to the nearest integer (pos_int) to align with the benchmark keys. It then retrieves the expected CTR using:

expected_ctr = self.EXPECTED_CTR.get(pos_int, 0.005)

This safeguards against out-of-range positions by defaulting to 0.5%. If the input data lacks a CTR field, the method automatically recomputes it from clicks / impressions before proceeding.

Threshold-Based Scoring Logic

The actual CTR is compared against four percentage thresholds of the expected value: 30%, 50%, 70%, and 90%. The scoring algorithm assigns discrete values based on where the actual performance falls:

  • Below 30% of expected → Score 100 (critical opportunity)
  • 30-50% of expected → Score 70
  • 50-70% of expected → Score 50
  • 70-90% of expected → Score 30
  • Above 90% of expected → Score 0 (performing at or above benchmark)

This inverse scoring system ensures that keywords with the largest CTR gaps receive the highest opportunity scores, directly signaling where metadata optimizations could drive immediate traffic increases.

Weighting and Integration in the Final Score

The raw ctr_score generated by _calculate_ctr_score feeds into the composite opportunity calculation within calculate_score (lines 144-152). Here, the CTR component receives a fixed weight of 5% (0.05) in the weighted sum alongside other SEO factors like difficulty and commercial intent.

This weighting ensures that CTR underperformance meaningfully influences the final priority ranking without overwhelming other signals. In QUICK_WIN opportunity analyses—where low-hanging fruit is prioritized—a low CTR score can significantly boost a keyword's overall priority despite moderate difficulty ratings.

Practical Implementation Example

The following example demonstrates how opportunity_scorer.py evaluates a keyword ranking at position 12 with subpar click-through performance:


# Example: scoring a keyword with a low CTR at position 12

from data_sources.modules.opportunity_scorer import OpportunityScorer, OpportunityType

keyword = {
    "position": 12,            # SERP rank

    "ctr": 0.006,              # 0.6% actual CTR (well below expected ~1.8%)

    "impressions": 2000,
    "clicks": 12,
    "commercial_intent": 1.2,
}
scorer = OpportunityScorer()
result = scorer.calculate_score(
    keyword_data=keyword,
    opportunity_type=OpportunityType.QUICK_WIN,
    difficulty=35,
)

print("CTR score :", result["score_breakdown"]["ctr_score"])   # → 100 (huge opportunity)

print("Overall   :", result["final_score"])

Adjusting the CTR closer to the benchmark reduces the opportunity score as the performance gap narrows:


# Adjust the CTR to see the effect on the CTR score

keyword["ctr"] = 0.015   # 1.5% (closer to expected 1.8%)

result = scorer.calculate_score(keyword_data=keyword,
                                opportunity_type=OpportunityType.QUICK_WIN,
                                difficulty=35)

print("CTR score :", result["score_breakdown"]["ctr_score"])   # → 70

Modules like research_quick_wins.py and research_competitor_gaps.py utilize this same scoring pipeline, demonstrating how position-specific CTR benchmarks provide consistent evaluation criteria across different SEO research workflows.

Summary

  • EXPECTED_CTR stores industry-average click-through rates for positions 1-20 at data_sources/modules/opportunity_scorer.py lines 37-59
  • _calculate_ctr_score (lines 74-104) compares actual CTR against 30%, 50%, 70%, and 90% thresholds of the position-specific benchmark
  • Scores range from 100 (severe underperformance) to 0 (meeting or exceeding expectations)
  • The CTR component carries a 5% weight in the final opportunity score calculated at lines 144-152
  • A 0.5% fallback CTR protects the calculation when processing keywords ranked beyond position 20

Frequently Asked Questions

What data structure stores the position-specific CTR benchmarks?

The benchmarks reside in the EXPECTED_CTR dictionary defined at the class level in opportunity_scorer.py. This structure maps integer SERP positions (1-20) to float values representing typical click-through rates, with rank 1 set to approximately 31.6% and rank 20 set to roughly 1.0%.

How does the scorer handle keywords with missing CTR data?

When the ctr key is absent from the input dictionary, _calculate_ctr_score automatically computes the metric by dividing clicks by impressions. If both are missing or zero, the calculation proceeds with available data, though the fallback mechanism for position lookup ensures the scoring never fails completely.

Why does the CTR score use an inverse scale where higher numbers indicate worse performance?

The 100-to-30 scoring range represents an "opportunity magnitude" rather than a quality grade. A score of 100 indicates the largest possible gap between expected and actual performance, signaling the highest potential ROI from optimization efforts like rewriting title tags or enhancing rich snippet eligibility.

What happens when a keyword ranks outside the top 20 positions?

For positions beyond 20, the lookup self.EXPECTED_CTR.get(pos_int, 0.005) returns a default value of 0.005 (0.5%). This conservative estimate prevents IndexError exceptions while ensuring that low-ranking keywords with measurable CTR still receive appropriate scoring relative to typical page-two performance.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →