How KeywordAnalyzer Detects Keyword Stuffing: Thresholds and Logic in SEO Machine
KeywordAnalyzer detects keyword stuffing using three concurrent checks—overall density caps at 2.5–3%, paragraph density caps at 5%, and consecutive sentence repetition caps at 3–5 sentences—with risk levels escalating from "low" to "high" based on severity.
The KeywordAnalyzer class in the TheCraigHewitt/seomachine repository evaluates content for SEO quality, including aggressive keyword stuffing detection. By analyzing density metrics at both the document and paragraph levels alongside repetition patterns, the analyzer flags over-optimization before it harms search rankings.
The Three-Pillar Detection Logic in _detect_keyword_stuffing
Inside data_sources/modules/keyword_analyzer.py, the private method _detect_keyword_stuffing (lines 312–365) implements a multi-layered validation system. This method returns a dictionary containing risk_level ("none", "low", "medium", or "high"), a list of warnings, and a boolean safe flag indicating whether the content passes inspection.
Overall Keyword Density Thresholds
The primary check evaluates the percentage of total words that match the target keyword. According to lines 324 and 327:
- Density > 3.0%: Triggers a "very high" density warning and sets
risk_levelto"high" - Density > 2.5%: Triggers a "high" density warning and sets
risk_levelto"medium"
These thresholds prevent the classic SEO mistake of repeating the primary keyword excessively throughout the entire document.
Paragraph-Level Density Validation
The analyzer scrutinizes each paragraph individually to catch localized stuffing that might evade document-wide averages. As implemented on lines 339–340:
- Paragraph density > 5%: Generates a specific warning for that paragraph
- Automatic escalation: Any paragraph exceeding 5% density immediately upgrades the overall
risk_levelto"high"regardless of the document's total density
This prevents writers from "hiding" high-density sections within otherwise balanced content.
Consecutive Sentence Repetition Analysis
The third check tracks how many back-to-back sentences contain the keyword, detecting aggressive repetition patterns. Based on lines 354 and 357:
- ≥ 5 consecutive sentences: Sets
risk_levelto"high" - ≥ 3 consecutive sentences: Sets
risk_levelto"low"
This catches robotic writing styles where the keyword appears in every sentence of a section.
Warning Thresholds at a Glance
| Check | Threshold | Risk Level | Source Line |
|---|---|---|---|
| Overall density | > 3.0% | High | 324 |
| Overall density | > 2.5% | Medium | 327 |
| Paragraph density | > 5.0% | High | 340 |
| Consecutive sentences | ≥ 5 | High | 354 |
| Consecutive sentences | ≥ 3 | Low | 357 |
Integration with the Analysis Pipeline
The stuffing detection integrates seamlessly into the main analysis workflow. During execution of the analyze method (lines 74–79), the system calculates the primary keyword's density and passes it to _detect_keyword_stuffing:
stuffing_result = self._detect_keyword_stuffing(
content,
primary_keyword,
primary_analysis['density']
)
The returned dictionary is stored under result['keyword_stuffing'] (line 100), making it available to downstream components. The recommendations generator (_generate_recommendations, lines 611–664) specifically checks keyword_stuffing['safe'] on line 62, injecting warning messages into the final report whenever the flag is False.
Practical Implementation Examples
Basic Usage with High-Risk Content
from data_sources.modules.keyword_analyzer import analyze_keywords
sample = """
# Start a Podcast
Starting a podcast is fun. Start a podcast today.
Start a podcast now! Start a podcast, start a podcast, start a podcast.
"""
result = analyze_keywords(sample, "start a podcast")
print(result["keyword_stuffing"])
Output:
{
"risk_level": "high",
"warnings": [
"Keyword density 4.2% is very high (over 3%)",
"Paragraph 1 has very high keyword density (7.5%)",
"Keyword appears in 5 consecutive sentences"
],
"safe": false
}
This example demonstrates how multiple threshold violations compound. The 4.2% overall density exceeds the 3% ceiling, the first paragraph hits 7.5% (above the 5% paragraph limit), and five consecutive sentences trigger the repetition alert.
Direct Method Invocation
from data_sources.modules.keyword_analyzer import KeywordAnalyzer
analyzer = KeywordAnalyzer()
content = "Keyword keyword keyword keyword keyword."
stuffing = analyzer._detect_keyword_stuffing(content, "keyword", 10.0)
print(stuffing)
Output:
{
"risk_level": "high",
"warnings": [
"Keyword density 10.0% is very high (over 3%)",
"Paragraph 1 has very high keyword density (100.0%)",
"Keyword appears in 5 consecutive sentences"
],
"safe": false
}
When calling _detect_keyword_stuffing directly, you pass the pre-calculated density (10.0% in this case) as the third argument, allowing the method to evaluate against the threshold matrix immediately.
Summary
- Three concurrent checks power the detection: document density (lines 324–327), paragraph density (lines 339–340), and consecutive sentence patterns (lines 354–357).
- Hard thresholds trigger at 2.5% (medium risk), 3.0% (high risk), and 5% paragraph density (immediate high risk).
- Sentence repetition escalates risk starting at 3 consecutive mentions and peaks at 5.
- Safe content requires
risk_levelof"none"or"low"to pass the booleansafecheck used by the recommendation engine. - Source location: All logic resides in
data_sources/modules/keyword_analyzer.pywithin theKeywordAnalyzerclass.
Frequently Asked Questions
What density percentage triggers a "high" risk warning in KeywordAnalyzer?
KeywordAnalyzer assigns "high" risk when overall keyword density exceeds 3.0%, when any paragraph exceeds 5.0%, or when the keyword appears in 5 or more consecutive sentences. These thresholds operate independently, meaning violating any single rule is sufficient to trigger the high-risk designation.
Is a 2% keyword density considered safe by the analyzer?
Yes. A 2% density falls below the 2.5% threshold that triggers the first warning level, so the analyzer would return "none" or "low" risk depending on other factors like paragraph distribution and sentence repetition. According to the logic on line 327, warnings only begin at densities strictly greater than 2.5%.
How does the paragraph-level check prevent "hidden" keyword stuffing?
By analyzing para_density individually for each paragraph (line 339), the detector catches high-concentration sections that might be diluted by longer, keyword-sparse sections elsewhere in the document. The 5% paragraph cap ensures local context remains readable even if the overall document average appears reasonable.
Can the analyzer detect semantic variations or synonyms as stuffing?
No. The _detect_keyword_stuffing method performs exact-string matching against the primary_keyword parameter passed to it. It does not evaluate semantic variations, LSI keywords, or synonyms unless they are explicitly included in the analysis call as separate keyword targets.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →