How the Content Analyzer Agent Chains Python Modules for Comprehensive SEO Analysis
The Content Analyzer agent orchestrates five specialized Python modules—Search Intent Analyzer, Keyword Analyzer, Content Length Comparator, Readability Scorer, and SEO Quality Rater—in a deterministic sequence to transform raw article drafts into data-rich, actionable SEO audits.
The Content Analyzer agent serves as the analytical engine within TheCraigHewitt/seomachine, executing a structured pipeline that evaluates content against search engine results page (SERP) competitors. By systematically chaining specific analysis modules located in data_sources/modules/, the agent generates comprehensive metrics that downstream optimization agents consume for content improvement.
The Three-Stage Pipeline Architecture
According to the agent specification in .claude/agents/content-analyzer.md, the Content Analyzer agent implements a three-stage workflow designed to process article inputs methodically. This architecture ensures that each analytical layer builds upon the previous, creating a cumulative assessment of SEO readiness.
Stage 1: Extracting Content Context
Before invoking any analysis modules, the agent gathers essential inputs from the article object. This extraction phase collects the article body, meta title and description, primary and secondary keywords, word count, internal and external link tallies, and any available SERP data such as DataForSEO results. These contextual elements form the parameter payload passed to each subsequent module in the chain.
Stage 2: Executing the Deterministic Module Chain
The core functionality resides in the sequential execution of five analysis classes, each exposing an analyze method. The agent imports these from concrete implementations in the data_sources/modules/ directory and invokes them in a strict order: Search Intent Analyzer → Keyword Analyzer → Content Length Comparator → Readability Scorer → SEO Quality Rater. This ordering is non-negotiable because downstream modules depend on intermediate outputs; specifically, the SEO Quality Rater requires the keyword density dictionary and link counts calculated by earlier steps.
SearchIntentAnalyzer (search_intent_analyzer.py) initiates the chain by classifying the primary keyword's intent. It uses keyword pattern matching, SERP feature mapping, and top-result content analysis to compute confidence scores. The agent calls:
intent_result = SearchIntentAnalyzer().analyze(
keyword=primary_keyword,
serp_features=serp_features,
top_results=top_results
)
KeywordAnalyzer (keyword_analyzer.py) follows, calculating keyword density, distribution, and semantic clustering. Leveraging TF-IDF, K-Means clustering, and custom stop-word handling, it surfaces primary and secondary keyword usage patterns. The agent invokes:
keyword_result = KeywordAnalyzer().analyze(
content=article_content,
primary_keyword=primary_keyword,
secondary_keywords=secondary_keywords,
target_density=1.5
)
ContentLengthComparator (content_length_comparator.py) benchmarks the article's word count against SERP competitors. Using requests and BeautifulSoup, it fetches top-ranking URLs, extracts word counts, and returns median and percentile statistics:
length_result = ContentLengthComparator().analyze(
keyword=primary_keyword,
your_word_count=word_count,
serp_results=serp_results,
fetch_content=True
)
ReadabilityScorer (readability_scorer.py) evaluates linguistic accessibility by wrapping the textstat library. It delivers Flesch Reading Ease scores, grade level assessments, sentence-length metrics, passive voice detection, and transition word analysis:
readability_result = ReadabilityScorer().analyze(content=article_content)
SeoQualityRater (seo_quality_rater.py) terminates the chain as the aggregation layer. It combines all previous signals with meta-element checks and link audits to emit a weighted SEO score (0-100) with category breakdowns. Crucially, it consumes the density output from KeywordAnalyzer:
seo_result = SeoQualityRater().analyze(
content=article_content,
meta_title=meta_title,
meta_description=meta_description,
primary_keyword=primary_keyword,
secondary_keywords=secondary_keywords,
keyword_density=keyword_result['primary_keyword']['density'],
internal_link_count=internal_links,
external_link_count=external_links
)
Stage 3: Synthesizing the Markdown Report
After all modules return Python dictionaries, the agent merges them into a richly structured Markdown report defined in the "Content Analysis Report" template within content-analyzer.md. This synthesis produces an executive summary, individual analytical sections for intent and readability, actionable recommendations prioritized by severity, and a competitive positioning snapshot that juxtaposes the article's metrics against SERP medians.
Implementation: Chaining Modules in Practice
The following orchestrator script demonstrates the exact import paths and execution order used by the Content Analyzer agent. This implementation mirrors the deterministic sequence specified in the agent documentation and handles the dependency injection required by the SEO Quality Rater.
from data_sources.modules.search_intent_analyzer import SearchIntentAnalyzer
from data_sources.modules.keyword_analyzer import KeywordAnalyzer
from data_sources.modules.content_length_comparator import ContentLengthComparator
from data_sources.modules.readability_scorer import ReadabilityScorer
from data_sources.modules.seo_quality_rater import SeoQualityRater
def run_content_analysis(article):
# Extract context from article dictionary
content = article["body"]
meta_title = article["meta_title"]
meta_description = article["meta_description"]
primary_kw = article["primary_keyword"]
secondary_kws = article.get("secondary_keywords", [])
word_count = len(content.split())
internal_links = article.get("internal_links", 0)
external_links = article.get("external_links", 0)
serp_features = article.get("serp_features")
top_results = article.get("top_results")
serp_results = article.get("serp_results")
# Execute module chain in prescribed order
intent = SearchIntentAnalyzer().analyze(
keyword=primary_kw,
serp_features=serp_features,
top_results=top_results,
)
kw = KeywordAnalyzer().analyze(
content=content,
primary_keyword=primary_kw,
secondary_keywords=secondary_kws,
target_density=1.5,
)
length = ContentLengthComparator().analyze(
keyword=primary_kw,
your_word_count=word_count,
serp_results=serp_results,
fetch_content=True,
)
readability = ReadabilityScorer().analyze(content=content)
# Final aggregation depends on previous outputs
seo = SeoQualityRater().analyze(
content=content,
meta_title=meta_title,
meta_description=meta_description,
primary_keyword=primary_kw,
secondary_keywords=secondary_kws,
keyword_density=kw["primary_keyword"]["density"],
internal_link_count=internal_links,
external_link_count=external_links,
)
return {
"intent": intent,
"keyword": kw,
"length": length,
"readability": readability,
"seo_score": seo,
}
Summary
- The Content Analyzer agent resides in TheCraigHewitt/seomachine repository and functions as the primary data-driven auditing component.
- Five specialized Python modules in
data_sources/modules/perform distinct analytical tasks: intent classification, keyword optimization, length benchmarking, readability scoring, and composite SEO rating. - The agent enforces a deterministic execution sequence—intent → keyword → length → readability → SEO rating—because the final SEO Quality Rater requires density metrics and link counts from preceding modules.
- Each module exposes a consistent
analyzeinterface, enabling the agent to invoke them with specific parameter payloads extracted from the article context. - Results synthesize into a structured Markdown report template defined in
.claude/agents/content-analyzer.md, providing actionable recommendations prioritized by severity.
Frequently Asked Questions
Why must the analysis modules execute in a specific order?
The SEO Quality Rater requires intermediate outputs generated by earlier modules in the chain. Specifically, it depends on the keyword density dictionary returned by KeywordAnalyzer and the internal/external link counts gathered during the initial context extraction phase. Executing modules out of sequence would result in missing dependencies and runtime errors when the final aggregation attempts to calculate the composite SEO score.
What Python libraries power the individual analysis modules?
According to the source implementations, KeywordAnalyzer leverages scikit-learn's K-Means clustering and TF-IDF vectorization alongside custom stop-word handling. ContentLengthComparator uses requests for HTTP fetching and BeautifulSoup for HTML parsing to extract competitor word counts. ReadabilityScorer wraps the textstat library to compute Flesch Reading Ease, grade levels, and sentence complexity metrics.
How does the agent handle missing SERP data during analysis?
The module signatures accept optional parameters such as serp_features, top_results, and serp_results. When these values are absent, the SearchIntentAnalyzer and ContentLengthComparator degrade gracefully, either returning intent classifications based solely on keyword pattern matching or skipping competitive length benchmarking. The agent stores these as optional dictionary keys using .get() methods, ensuring the pipeline continues execution without raising KeyError exceptions.
Where can I modify the SEO scoring weights or add new analysis modules?
The SeoQualityRater implementation in data_sources/modules/seo_quality_rater.py contains the weighted scoring logic that produces the 0-100 composite score. To adjust category weights or add new signals, modify this file's analysis method. For adding entirely new analysis dimensions, create a new module in data_sources/modules/ exposing an analyze method, then insert it into the execution chain within the agent's orchestration logic before the final SeoQualityRater invocation.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →