How seo_quality_rater.py Calculates the 0‑100 SEO Score: Inside the Rule‑Based Methodology

The 0‑100 SEO score in seo_quality_rater.py is a weighted aggregate of six independent category scores, each calculated by deducting penalty points from a baseline of 100 based on configurable thresholds for word count, keyword density, meta tags, headings, links, and readability.

The seo_quality_rater.py module in the TheCraigHewitt/seomachine repository implements a deterministic, rule‑based scoring engine that evaluates content against SEO best practices. This methodology breaks down complex content analysis into discrete, measurable components, culminating in a normalized 0‑100 Overall SEO Score (O‑IOO) that reflects publishing readiness.

The Three‑Layer Scoring Architecture

The engine operates through three distinct phases defined in data_sources/modules/seo_quality_rater.py, separating data extraction from evaluation and final aggregation.

Layer 1: Structure Analysis (_analyze_structure)

The rate() method first invokes _analyze_structure() (source lines 78‑80) to parse the raw content and extract quantitative metrics:

  • H1, H2, and H3 counts and text content
  • Total word_count
  • Average paragraph length (avg_paragraph_length)
  • Boolean flags for keyword placement: keyword_in_h1, keyword_in_first_100, and h2_with_keyword

These structured data points serve as the input for all subsequent penalty calculations.

Layer 2: Category Scoring

Six independent sub‑scorers evaluate specific SEO dimensions. Each begins with a perfect score of 100 and applies deductions based on thresholds defined in self.guidelines.

Scorer Purpose Source Location
_score_content Word count, paragraph length lines 21‑58
_score_keyword_optimization Keyword placement (H1, first 100 words, H2), density, secondary coverage lines 60‑84
_score_meta_elements Meta title/description length and keyword inclusion lines 86‑101
_score_structure H1 presence/uniqueness, H2 count lines 103‑115
_score_links Internal/external link counts vs. min/optimal thresholds lines 117‑133
_score_readability Average sentence length, proportion of very long sentences, presence of lists lines 135‑152

Layer 3: Weighted Aggregation

After the six category scores are computed, rate() applies a fixed weight map (source lines 104‑110) to derive the final 0‑100 Overall SEO Score (O‑IOO):

weights = {
    'content': 0.20,
    'keywords': 0.25,
    'meta': 0.15,
    'structure': 0.15,
    'links': 0.15,
    'readability': 0.10,
}

The weighted sum is calculated as follows (source lines 112‑121):

overall_score = (
    content_score['score'] * weights['content'] +
    keyword_score['score'] * weights['keywords'] +
    meta_score['score'] * weights['meta'] +
    structure_score['score'] * weights['structure'] +
    link_score['score'] * weights['links'] +
    readability_score['score'] * weights['readability']
)

The result is rounded to one decimal place. A letter grade (A‑F) is then derived via _get_grade() (lines 124‑131).

Configurable SEO Guidelines and Thresholds

The scoring engine relies on a dictionary of thresholds instantiated in _default_guidelines() (source lines 24‑49). These values define the boundaries for penalties and bonuses:

Area Metric Thresholds
Word count min / optimal / max 2000 / 2500 / 3000
Primary keyword density min / max 1 % / 2 %
Internal links min / optimal 3 / 5
External links min / optimal 2 / 3
Meta title length target range 50‑60 characters
Meta description length target range 150‑160 characters
H2 sections min / optimal 4 / 6
H2‑with‑keyword ratio target 0.33
Max sentence length threshold 25 words
Target reading level grade range 8‑10
Paragraph density sentences per paragraph 2‑4

Users can override these defaults by passing a custom guidelines dictionary to the SEOQualityRater constructor or the rate_seo_quality() helper.

Practical Implementation Example

The following example demonstrates how to invoke the scoring engine using the public rate_seo_quality() wrapper (source lines 152‑162):

from data_sources.modules.seo_quality_rater import rate_seo_quality

sample = """

# How to Start a Podcast

Starting a podcast is easier than you think. This complete guide shows you how to start a podcast from scratch.

## Choose Your Topic

Pick a topic you're passionate about. Your podcast topic should resonate with your target audience.

## Get Equipment

You'll need a microphone, headphones, and recording software.

## Record Your First Episode

Start recording! Don't worry about perfection on your first try.

## Publish Your Podcast

Upload to a podcast hosting platform and distribute to directories.
"""

result = rate_seo_quality(
    content=sample,
    meta_title="How to Start a Podcast: Complete Guide for 2024",
    meta_description="Learn how to start a podcast from scratch with this step‑by‑step guide. Everything you need to know about equipment, recording, and publishing.",
    primary_keyword="start a podcast",
    secondary_keywords=["podcast hosting", "recording software"],
    keyword_density=1.8,
    internal_link_count=4,
    external_link_count=2,
)

print("Overall SEO score:", result["overall_score"])
print("Grade:", result["grade"])
print("Publishing ready?", result["publishing_ready"])

The function returns a dictionary containing the overall_score, grade, category_scores, critical_issues, warnings, suggestions, and a publishing_ready boolean (True when the score is ≥ 80 and no critical issues exist).

Summary

  • seo_quality_rater.py implements a deterministic, rule‑based scoring engine that evaluates content against configurable SEO best practices.
  • The 0‑100 Overall SEO Score is a weighted aggregate of six category scores: Content (20 %), Keywords (25 %), Meta (15 %), Structure (15 %), Links (15 %), and Readability (10 %).
  • Each category scorer starts at 100 and applies deductions based on thresholds defined in _default_guidelines(), covering word count (2000‑3000), keyword density (1‑2 %), internal/external link counts, and readability metrics.
  • The final payload includes granular feedback (critical issues, warnings, suggestions) and a publishing_ready flag triggered at scores ≥ 80 with zero critical issues.

Frequently Asked Questions

How does seo_quality_rater.py handle missing meta titles or descriptions?

The _score_meta_elements method (lines 86‑101) checks for the presence and character length of meta titles and descriptions. If these elements are missing or fall outside the 50‑60 character (title) or 150‑160 character (description) ranges defined in the guidelines, the scorer applies proportional penalties to the base score of 100, recording absent metadata as critical issues.

Can I customize the weights for the 0‑100 SEO score calculation?

Currently, the weight map is hardcoded in the rate() method (lines 104‑110) with fixed values emphasizing keyword optimization (25 %) and content depth (20 %). While the guidelines dictionary is fully configurable via the class constructor, adjusting the relative importance of readability, links, or meta elements requires modifying the weights dictionary directly in data_sources/modules/seo_quality_rater.py.

What is the difference between a critical issue and a warning in the scoring output?

Each category scorer maintains separate lists for critical issues, warnings, and suggestions. Critical issues represent severe SEO defects—such as a missing H1 tag, absent primary keyword in the first 100 words, or zero internal links—that trigger heavy score deductions (often 20‑40 points) and automatically disqualify the content from receiving a publishing_ready status. Warnings indicate suboptimal but acceptable conditions (e.g., keyword density slightly below 1 %), incurring smaller penalties (5‑10 points), while suggestions provide zero‑penalty recommendations for further optimization.

How does the scoring engine determine if content is "publishing ready"?

The publishing_ready boolean is set to True only when the overall_score is greater than or equal to 80 and the critical_issues list is empty. This logic, implemented in the rate() method, ensures that content meeting the weighted 0‑100 threshold but suffering from fundamental structural flaws (like missing primary keywords in headers) cannot pass the final quality gate, maintaining strict adherence to the SEO guidelines defined in data_sources/modules/seo_quality_rater.py.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →