# How KeywordAnalyzer Detects Keyword Stuffing: Thresholds and Logic in SEO Machine

> Learn how KeywordAnalyzer detects keyword stuffing with its density and repetition checks. Understand the thresholds and risk levels for better SEO.

- Repository: [Craig/seomachine](https://github.com/TheCraigHewitt/seomachine)
- Tags: deep-dive
- Published: 2026-03-12

---

**KeywordAnalyzer detects keyword stuffing using three concurrent checks—overall density caps at 2.5–3%, paragraph density caps at 5%, and consecutive sentence repetition caps at 3–5 sentences—with risk levels escalating from "low" to "high" based on severity.**

The `KeywordAnalyzer` class in the [TheCraigHewitt/seomachine](https://github.com/TheCraigHewitt/seomachine) repository evaluates content for SEO quality, including aggressive keyword stuffing detection. By analyzing density metrics at both the document and paragraph levels alongside repetition patterns, the analyzer flags over-optimization before it harms search rankings.

## The Three-Pillar Detection Logic in `_detect_keyword_stuffing`

Inside [`data_sources/modules/keyword_analyzer.py`](https://github.com/TheCraigHewitt/seomachine/blob/main/data_sources/modules/keyword_analyzer.py), the private method `_detect_keyword_stuffing` (lines 312–365) implements a multi-layered validation system. This method returns a dictionary containing `risk_level` (`"none"`, `"low"`, `"medium"`, or `"high"`), a list of `warnings`, and a boolean `safe` flag indicating whether the content passes inspection.

### Overall Keyword Density Thresholds

The primary check evaluates the percentage of total words that match the target keyword. According to lines 324 and 327:

- **Density > 3.0%**: Triggers a "very high" density warning and sets `risk_level` to `"high"`
- **Density > 2.5%**: Triggers a "high" density warning and sets `risk_level` to `"medium"`

These thresholds prevent the classic SEO mistake of repeating the primary keyword excessively throughout the entire document.

### Paragraph-Level Density Validation

The analyzer scrutinizes each paragraph individually to catch localized stuffing that might evade document-wide averages. As implemented on lines 339–340:

- **Paragraph density > 5%**: Generates a specific warning for that paragraph
- **Automatic escalation**: Any paragraph exceeding 5% density immediately upgrades the overall `risk_level` to `"high"` regardless of the document's total density

This prevents writers from "hiding" high-density sections within otherwise balanced content.

### Consecutive Sentence Repetition Analysis

The third check tracks how many back-to-back sentences contain the keyword, detecting aggressive repetition patterns. Based on lines 354 and 357:

- **≥ 5 consecutive sentences**: Sets `risk_level` to `"high"`
- **≥ 3 consecutive sentences**: Sets `risk_level` to `"low"`

This catches robotic writing styles where the keyword appears in every sentence of a section.

## Warning Thresholds at a Glance

| Check | Threshold | Risk Level | Source Line |
|-------|-----------|------------|-------------|
| **Overall density** | > 3.0% | High | 324 |
| **Overall density** | > 2.5% | Medium | 327 |
| **Paragraph density** | > 5.0% | High | 340 |
| **Consecutive sentences** | ≥ 5 | High | 354 |
| **Consecutive sentences** | ≥ 3 | Low | 357 |

## Integration with the Analysis Pipeline

The stuffing detection integrates seamlessly into the main analysis workflow. During execution of the `analyze` method (lines 74–79), the system calculates the primary keyword's density and passes it to `_detect_keyword_stuffing`:

```python
stuffing_result = self._detect_keyword_stuffing(
    content, 
    primary_keyword, 
    primary_analysis['density']
)

```

The returned dictionary is stored under `result['keyword_stuffing']` (line 100), making it available to downstream components. The recommendations generator (`_generate_recommendations`, lines 611–664) specifically checks `keyword_stuffing['safe']` on line 62, injecting warning messages into the final report whenever the flag is `False`.

## Practical Implementation Examples

### Basic Usage with High-Risk Content

```python
from data_sources.modules.keyword_analyzer import analyze_keywords

sample = """

# Start a Podcast

Starting a podcast is fun. Start a podcast today.
Start a podcast now! Start a podcast, start a podcast, start a podcast.
"""

result = analyze_keywords(sample, "start a podcast")
print(result["keyword_stuffing"])

```

**Output:**

```json
{
  "risk_level": "high",
  "warnings": [
    "Keyword density 4.2% is very high (over 3%)",
    "Paragraph 1 has very high keyword density (7.5%)",
    "Keyword appears in 5 consecutive sentences"
  ],
  "safe": false
}

```

This example demonstrates how multiple threshold violations compound. The 4.2% overall density exceeds the 3% ceiling, the first paragraph hits 7.5% (above the 5% paragraph limit), and five consecutive sentences trigger the repetition alert.

### Direct Method Invocation

```python
from data_sources.modules.keyword_analyzer import KeywordAnalyzer

analyzer = KeywordAnalyzer()
content = "Keyword keyword keyword keyword keyword."
stuffing = analyzer._detect_keyword_stuffing(content, "keyword", 10.0)
print(stuffing)

```

**Output:**

```json
{
  "risk_level": "high",
  "warnings": [
    "Keyword density 10.0% is very high (over 3%)",
    "Paragraph 1 has very high keyword density (100.0%)",
    "Keyword appears in 5 consecutive sentences"
  ],
  "safe": false
}

```

When calling `_detect_keyword_stuffing` directly, you pass the pre-calculated density (10.0% in this case) as the third argument, allowing the method to evaluate against the threshold matrix immediately.

## Summary

- **Three concurrent checks** power the detection: document density (lines 324–327), paragraph density (lines 339–340), and consecutive sentence patterns (lines 354–357).
- **Hard thresholds** trigger at 2.5% (medium risk), 3.0% (high risk), and 5% paragraph density (immediate high risk).
- **Sentence repetition** escalates risk starting at 3 consecutive mentions and peaks at 5.
- **Safe content** requires `risk_level` of `"none"` or `"low"` to pass the boolean `safe` check used by the recommendation engine.
- **Source location**: All logic resides in [`data_sources/modules/keyword_analyzer.py`](https://github.com/TheCraigHewitt/seomachine/blob/main/data_sources/modules/keyword_analyzer.py) within the `KeywordAnalyzer` class.

## Frequently Asked Questions

### What density percentage triggers a "high" risk warning in KeywordAnalyzer?

KeywordAnalyzer assigns `"high"` risk when overall keyword density exceeds **3.0%**, when any paragraph exceeds **5.0%**, or when the keyword appears in **5 or more consecutive sentences**. These thresholds operate independently, meaning violating any single rule is sufficient to trigger the high-risk designation.

### Is a 2% keyword density considered safe by the analyzer?

Yes. A 2% density falls below the **2.5% threshold** that triggers the first warning level, so the analyzer would return `"none"` or `"low"` risk depending on other factors like paragraph distribution and sentence repetition. According to the logic on line 327, warnings only begin at densities strictly greater than 2.5%.

### How does the paragraph-level check prevent "hidden" keyword stuffing?

By analyzing `para_density` individually for each paragraph (line 339), the detector catches high-concentration sections that might be diluted by longer, keyword-sparse sections elsewhere in the document. The **5% paragraph cap** ensures local context remains readable even if the overall document average appears reasonable.

### Can the analyzer detect semantic variations or synonyms as stuffing?

No. The `_detect_keyword_stuffing` method performs exact-string matching against the `primary_keyword` parameter passed to it. It does not evaluate semantic variations, LSI keywords, or synonyms unless they are explicitly included in the analysis call as separate keyword targets.