How seo_quality_rater.py Calculates the 0‑100 SEO Score: Inside the Rule‑Based Methodology
The 0‑100 SEO score in seo_quality_rater.py is a weighted aggregate of six independent category scores, each calculated by deducting penalty points from a baseline of 100 based on configurable thresholds for word count, keyword density, meta tags, headings, links, and readability.
The seo_quality_rater.py module in the TheCraigHewitt/seomachine repository implements a deterministic, rule‑based scoring engine that evaluates content against SEO best practices. This methodology breaks down complex content analysis into discrete, measurable components, culminating in a normalized 0‑100 Overall SEO Score (O‑IOO) that reflects publishing readiness.
The Three‑Layer Scoring Architecture
The engine operates through three distinct phases defined in data_sources/modules/seo_quality_rater.py, separating data extraction from evaluation and final aggregation.
Layer 1: Structure Analysis (_analyze_structure)
The rate() method first invokes _analyze_structure() (source lines 78‑80) to parse the raw content and extract quantitative metrics:
- H1, H2, and H3 counts and text content
- Total
word_count - Average paragraph length (
avg_paragraph_length) - Boolean flags for keyword placement:
keyword_in_h1,keyword_in_first_100, andh2_with_keyword
These structured data points serve as the input for all subsequent penalty calculations.
Layer 2: Category Scoring
Six independent sub‑scorers evaluate specific SEO dimensions. Each begins with a perfect score of 100 and applies deductions based on thresholds defined in self.guidelines.
| Scorer | Purpose | Source Location |
|---|---|---|
_score_content |
Word count, paragraph length | lines 21‑58 |
_score_keyword_optimization |
Keyword placement (H1, first 100 words, H2), density, secondary coverage | lines 60‑84 |
_score_meta_elements |
Meta title/description length and keyword inclusion | lines 86‑101 |
_score_structure |
H1 presence/uniqueness, H2 count | lines 103‑115 |
_score_links |
Internal/external link counts vs. min/optimal thresholds | lines 117‑133 |
_score_readability |
Average sentence length, proportion of very long sentences, presence of lists | lines 135‑152 |
Layer 3: Weighted Aggregation
After the six category scores are computed, rate() applies a fixed weight map (source lines 104‑110) to derive the final 0‑100 Overall SEO Score (O‑IOO):
weights = {
'content': 0.20,
'keywords': 0.25,
'meta': 0.15,
'structure': 0.15,
'links': 0.15,
'readability': 0.10,
}
The weighted sum is calculated as follows (source lines 112‑121):
overall_score = (
content_score['score'] * weights['content'] +
keyword_score['score'] * weights['keywords'] +
meta_score['score'] * weights['meta'] +
structure_score['score'] * weights['structure'] +
link_score['score'] * weights['links'] +
readability_score['score'] * weights['readability']
)
The result is rounded to one decimal place. A letter grade (A‑F) is then derived via _get_grade() (lines 124‑131).
Configurable SEO Guidelines and Thresholds
The scoring engine relies on a dictionary of thresholds instantiated in _default_guidelines() (source lines 24‑49). These values define the boundaries for penalties and bonuses:
| Area | Metric | Thresholds |
|---|---|---|
| Word count | min / optimal / max | 2000 / 2500 / 3000 |
| Primary keyword density | min / max | 1 % / 2 % |
| Internal links | min / optimal | 3 / 5 |
| External links | min / optimal | 2 / 3 |
| Meta title length | target range | 50‑60 characters |
| Meta description length | target range | 150‑160 characters |
| H2 sections | min / optimal | 4 / 6 |
| H2‑with‑keyword ratio | target | 0.33 |
| Max sentence length | threshold | 25 words |
| Target reading level | grade range | 8‑10 |
| Paragraph density | sentences per paragraph | 2‑4 |
Users can override these defaults by passing a custom guidelines dictionary to the SEOQualityRater constructor or the rate_seo_quality() helper.
Practical Implementation Example
The following example demonstrates how to invoke the scoring engine using the public rate_seo_quality() wrapper (source lines 152‑162):
from data_sources.modules.seo_quality_rater import rate_seo_quality
sample = """
# How to Start a Podcast
Starting a podcast is easier than you think. This complete guide shows you how to start a podcast from scratch.
## Choose Your Topic
Pick a topic you're passionate about. Your podcast topic should resonate with your target audience.
## Get Equipment
You'll need a microphone, headphones, and recording software.
## Record Your First Episode
Start recording! Don't worry about perfection on your first try.
## Publish Your Podcast
Upload to a podcast hosting platform and distribute to directories.
"""
result = rate_seo_quality(
content=sample,
meta_title="How to Start a Podcast: Complete Guide for 2024",
meta_description="Learn how to start a podcast from scratch with this step‑by‑step guide. Everything you need to know about equipment, recording, and publishing.",
primary_keyword="start a podcast",
secondary_keywords=["podcast hosting", "recording software"],
keyword_density=1.8,
internal_link_count=4,
external_link_count=2,
)
print("Overall SEO score:", result["overall_score"])
print("Grade:", result["grade"])
print("Publishing ready?", result["publishing_ready"])
The function returns a dictionary containing the overall_score, grade, category_scores, critical_issues, warnings, suggestions, and a publishing_ready boolean (True when the score is ≥ 80 and no critical issues exist).
Summary
- seo_quality_rater.py implements a deterministic, rule‑based scoring engine that evaluates content against configurable SEO best practices.
- The 0‑100 Overall SEO Score is a weighted aggregate of six category scores: Content (20 %), Keywords (25 %), Meta (15 %), Structure (15 %), Links (15 %), and Readability (10 %).
- Each category scorer starts at 100 and applies deductions based on thresholds defined in
_default_guidelines(), covering word count (2000‑3000), keyword density (1‑2 %), internal/external link counts, and readability metrics. - The final payload includes granular feedback (critical issues, warnings, suggestions) and a
publishing_readyflag triggered at scores ≥ 80 with zero critical issues.
Frequently Asked Questions
How does seo_quality_rater.py handle missing meta titles or descriptions?
The _score_meta_elements method (lines 86‑101) checks for the presence and character length of meta titles and descriptions. If these elements are missing or fall outside the 50‑60 character (title) or 150‑160 character (description) ranges defined in the guidelines, the scorer applies proportional penalties to the base score of 100, recording absent metadata as critical issues.
Can I customize the weights for the 0‑100 SEO score calculation?
Currently, the weight map is hardcoded in the rate() method (lines 104‑110) with fixed values emphasizing keyword optimization (25 %) and content depth (20 %). While the guidelines dictionary is fully configurable via the class constructor, adjusting the relative importance of readability, links, or meta elements requires modifying the weights dictionary directly in data_sources/modules/seo_quality_rater.py.
What is the difference between a critical issue and a warning in the scoring output?
Each category scorer maintains separate lists for critical issues, warnings, and suggestions. Critical issues represent severe SEO defects—such as a missing H1 tag, absent primary keyword in the first 100 words, or zero internal links—that trigger heavy score deductions (often 20‑40 points) and automatically disqualify the content from receiving a publishing_ready status. Warnings indicate suboptimal but acceptable conditions (e.g., keyword density slightly below 1 %), incurring smaller penalties (5‑10 points), while suggestions provide zero‑penalty recommendations for further optimization.
How does the scoring engine determine if content is "publishing ready"?
The publishing_ready boolean is set to True only when the overall_score is greater than or equal to 80 and the critical_issues list is empty. This logic, implemented in the rate() method, ensures that content meeting the weighted 0‑100 threshold but suffering from fundamental structural flaws (like missing primary keywords in headers) cannot pass the final quality gate, maintaining strict adherence to the SEO guidelines defined in data_sources/modules/seo_quality_rater.py.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →