# How seo_quality_rater.py Calculates the 0‑100 SEO Score: Inside the Rule‑Based Methodology

> Understand the seo_quality_rater.py methodology. Discover how its rule-based system calculates a 0-100 SEO score using word count, keywords, meta tags, headings, links, and readability.

- Repository: [Craig/seomachine](https://github.com/TheCraigHewitt/seomachine)
- Tags: deep-dive
- Published: 2026-03-12

---

**The 0‑100 SEO score in [`seo_quality_rater.py`](https://github.com/TheCraigHewitt/seomachine/blob/main/seo_quality_rater.py) is a weighted aggregate of six independent category scores, each calculated by deducting penalty points from a baseline of 100 based on configurable thresholds for word count, keyword density, meta tags, headings, links, and readability.**

The [`seo_quality_rater.py`](https://github.com/TheCraigHewitt/seomachine/blob/main/seo_quality_rater.py) module in the [TheCraigHewitt/seomachine](https://github.com/TheCraigHewitt/seomachine) repository implements a deterministic, rule‑based scoring engine that evaluates content against SEO best practices. This methodology breaks down complex content analysis into discrete, measurable components, culminating in a normalized 0‑100 Overall SEO Score (O‑IOO) that reflects publishing readiness.

## The Three‑Layer Scoring Architecture

The engine operates through three distinct phases defined in [`data_sources/modules/seo_quality_rater.py`](https://github.com/TheCraigHewitt/seomachine/blob/main/data_sources/modules/seo_quality_rater.py), separating data extraction from evaluation and final aggregation.

### Layer 1: Structure Analysis (`_analyze_structure`)

The `rate()` method first invokes `_analyze_structure()` ([source lines 78‑80](https://github.com/TheCraigHewitt/seomachine/blob/main/data_sources/modules/seo_quality_rater.py#L78-L80)) to parse the raw content and extract quantitative metrics:

- H1, H2, and H3 counts and text content
- Total `word_count`
- Average paragraph length (`avg_paragraph_length`)
- Boolean flags for keyword placement: `keyword_in_h1`, `keyword_in_first_100`, and `h2_with_keyword`

These structured data points serve as the input for all subsequent penalty calculations.

### Layer 2: Category Scoring

Six independent sub‑scorers evaluate specific SEO dimensions. Each begins with a perfect score of 100 and applies deductions based on thresholds defined in `self.guidelines`.

| Scorer | Purpose | Source Location |
|--------|---------|-----------------|
| `_score_content` | Word count, paragraph length | [lines 21‑58](https://github.com/TheCraigHewitt/seomachine/blob/main/data_sources/modules/seo_quality_rater.py#L21-L58) |
| `_score_keyword_optimization` | Keyword placement (H1, first 100 words, H2), density, secondary coverage | [lines 60‑84](https://github.com/TheCraigHewitt/seomachine/blob/main/data_sources/modules/seo_quality_rater.py#L60-L84) |
| `_score_meta_elements` | Meta title/description length and keyword inclusion | [lines 86‑101](https://github.com/TheCraigHewitt/seomachine/blob/main/data_sources/modules/seo_quality_rater.py#L86-L101) |
| `_score_structure` | H1 presence/uniqueness, H2 count | [lines 103‑115](https://github.com/TheCraigHewitt/seomachine/blob/main/data_sources/modules/seo_quality_rater.py#L103-L115) |
| `_score_links` | Internal/external link counts vs. min/optimal thresholds | [lines 117‑133](https://github.com/TheCraigHewitt/seomachine/blob/main/data_sources/modules/seo_quality_rater.py#L117-L133) |
| `_score_readability` | Average sentence length, proportion of very long sentences, presence of lists | [lines 135‑152](https://github.com/TheCraigHewitt/seomachine/blob/main/data_sources/modules/seo_quality_rater.py#L135-L152) |

### Layer 3: Weighted Aggregation

After the six category scores are computed, `rate()` applies a fixed weight map ([source lines 104‑110](https://github.com/TheCraigHewitt/seomachine/blob/main/data_sources/modules/seo_quality_rater.py#L104-L110)) to derive the final 0‑100 **Overall SEO Score** (O‑IOO):

```python
weights = {
    'content': 0.20,
    'keywords': 0.25,
    'meta': 0.15,
    'structure': 0.15,
    'links': 0.15,
    'readability': 0.10,
}

```

The weighted sum is calculated as follows ([source lines 112‑121](https://github.com/TheCraigHewitt/seomachine/blob/main/data_sources/modules/seo_quality_rater.py#L112-L121)):

```python
overall_score = (
    content_score['score'] * weights['content'] +
    keyword_score['score'] * weights['keywords'] +
    meta_score['score'] * weights['meta'] +
    structure_score['score'] * weights['structure'] +
    link_score['score'] * weights['links'] +
    readability_score['score'] * weights['readability']
)

```

The result is rounded to one decimal place. A letter grade (A‑F) is then derived via `_get_grade()` ([lines 124‑131](https://github.com/TheCraigHewitt/seomachine/blob/main/data_sources/modules/seo_quality_rater.py#L124-L131)).

## Configurable SEO Guidelines and Thresholds

The scoring engine relies on a dictionary of thresholds instantiated in `_default_guidelines()` ([source lines 24‑49](https://github.com/TheCraigHewitt/seomachine/blob/main/data_sources/modules/seo_quality_rater.py#L24-L49)). These values define the boundaries for penalties and bonuses:

| Area | Metric | Thresholds |
|------|--------|------------|
| **Word count** | min / optimal / max | 2000 / 2500 / 3000 |
| **Primary keyword density** | min / max | 1 % / 2 % |
| **Internal links** | min / optimal | 3 / 5 |
| **External links** | min / optimal | 2 / 3 |
| **Meta title length** | target range | 50‑60 characters |
| **Meta description length** | target range | 150‑160 characters |
| **H2 sections** | min / optimal | 4 / 6 |
| **H2‑with‑keyword ratio** | target | 0.33 |
| **Max sentence length** | threshold | 25 words |
| **Target reading level** | grade range | 8‑10 |
| **Paragraph density** | sentences per paragraph | 2‑4 |

Users can override these defaults by passing a custom `guidelines` dictionary to the `SEOQualityRater` constructor or the `rate_seo_quality()` helper.

## Practical Implementation Example

The following example demonstrates how to invoke the scoring engine using the public `rate_seo_quality()` wrapper ([source lines 152‑162](https://github.com/TheCraigHewitt/seomachine/blob/main/data_sources/modules/seo_quality_rater.py#L152-L162)):

```python
from data_sources.modules.seo_quality_rater import rate_seo_quality

sample = """

# How to Start a Podcast

Starting a podcast is easier than you think. This complete guide shows you how to start a podcast from scratch.

## Choose Your Topic

Pick a topic you're passionate about. Your podcast topic should resonate with your target audience.

## Get Equipment

You'll need a microphone, headphones, and recording software.

## Record Your First Episode

Start recording! Don't worry about perfection on your first try.

## Publish Your Podcast

Upload to a podcast hosting platform and distribute to directories.
"""

result = rate_seo_quality(
    content=sample,
    meta_title="How to Start a Podcast: Complete Guide for 2024",
    meta_description="Learn how to start a podcast from scratch with this step‑by‑step guide. Everything you need to know about equipment, recording, and publishing.",
    primary_keyword="start a podcast",
    secondary_keywords=["podcast hosting", "recording software"],
    keyword_density=1.8,
    internal_link_count=4,
    external_link_count=2,
)

print("Overall SEO score:", result["overall_score"])
print("Grade:", result["grade"])
print("Publishing ready?", result["publishing_ready"])

```

The function returns a dictionary containing the `overall_score`, `grade`, `category_scores`, `critical_issues`, `warnings`, `suggestions`, and a `publishing_ready` boolean (True when the score is ≥ 80 and no critical issues exist).

## Summary

- **seo_quality_rater.py** implements a deterministic, rule‑based scoring engine that evaluates content against configurable SEO best practices.
- The **0‑100 Overall SEO Score** is a weighted aggregate of six category scores: **Content** (20 %), **Keywords** (25 %), **Meta** (15 %), **Structure** (15 %), **Links** (15 %), and **Readability** (10 %).
- Each category scorer starts at 100 and applies deductions based on thresholds defined in `_default_guidelines()`, covering word count (2000‑3000), keyword density (1‑2 %), internal/external link counts, and readability metrics.
- The final payload includes granular feedback (critical issues, warnings, suggestions) and a `publishing_ready` flag triggered at scores ≥ 80 with zero critical issues.

## Frequently Asked Questions

### How does seo_quality_rater.py handle missing meta titles or descriptions?

The `_score_meta_elements` method ([lines 86‑101](https://github.com/TheCraigHewitt/seomachine/blob/main/data_sources/modules/seo_quality_rater.py#L86-L101)) checks for the presence and character length of meta titles and descriptions. If these elements are missing or fall outside the 50‑60 character (title) or 150‑160 character (description) ranges defined in the guidelines, the scorer applies proportional penalties to the base score of 100, recording absent metadata as critical issues.

### Can I customize the weights for the 0‑100 SEO score calculation?

Currently, the weight map is hardcoded in the `rate()` method ([lines 104‑110](https://github.com/TheCraigHewitt/seomachine/blob/main/data_sources/modules/seo_quality_rater.py#L104-L110)) with fixed values emphasizing keyword optimization (25 %) and content depth (20 %). While the `guidelines` dictionary is fully configurable via the class constructor, adjusting the relative importance of readability, links, or meta elements requires modifying the `weights` dictionary directly in [`data_sources/modules/seo_quality_rater.py`](https://github.com/TheCraigHewitt/seomachine/blob/main/data_sources/modules/seo_quality_rater.py).

### What is the difference between a critical issue and a warning in the scoring output?

Each category scorer maintains separate lists for **critical issues**, **warnings**, and **suggestions**. Critical issues represent severe SEO defects—such as a missing H1 tag, absent primary keyword in the first 100 words, or zero internal links—that trigger heavy score deductions (often 20‑40 points) and automatically disqualify the content from receiving a `publishing_ready` status. Warnings indicate suboptimal but acceptable conditions (e.g., keyword density slightly below 1 %), incurring smaller penalties (5‑10 points), while suggestions provide zero‑penalty recommendations for further optimization.

### How does the scoring engine determine if content is "publishing ready"?

The `publishing_ready` boolean is set to `True` only when the `overall_score` is greater than or equal to 80 **and** the `critical_issues` list is empty. This logic, implemented in the `rate()` method, ensures that content meeting the weighted 0‑100 threshold but suffering from fundamental structural flaws (like missing primary keywords in headers) cannot pass the final quality gate, maintaining strict adherence to the SEO guidelines defined in [`data_sources/modules/seo_quality_rater.py`](https://github.com/TheCraigHewitt/seomachine/blob/main/data_sources/modules/seo_quality_rater.py).