# How the Content Analyzer Agent Chains Python Modules for Comprehensive SEO Analysis

> Discover how the Content Analyzer agent chains five Python modules to deliver comprehensive SEO analysis. Transform drafts into actionable SEO audits with expert insights.

- Repository: [Craig/seomachine](https://github.com/TheCraigHewitt/seomachine)
- Tags: internals
- Published: 2026-03-12

---

**The Content Analyzer agent orchestrates five specialized Python modules—Search Intent Analyzer, Keyword Analyzer, Content Length Comparator, Readability Scorer, and SEO Quality Rater—in a deterministic sequence to transform raw article drafts into data-rich, actionable SEO audits.**

The Content Analyzer agent serves as the analytical engine within TheCraigHewitt/seomachine, executing a structured pipeline that evaluates content against search engine results page (SERP) competitors. By systematically chaining specific analysis modules located in `data_sources/modules/`, the agent generates comprehensive metrics that downstream optimization agents consume for content improvement.

## The Three-Stage Pipeline Architecture

According to the agent specification in [`.claude/agents/content-analyzer.md`](https://github.com/TheCraigHewitt/seomachine/blob/main/.claude/agents/content-analyzer.md), the Content Analyzer agent implements a three-stage workflow designed to process article inputs methodically. This architecture ensures that each analytical layer builds upon the previous, creating a cumulative assessment of SEO readiness.

### Stage 1: Extracting Content Context

Before invoking any analysis modules, the agent gathers essential inputs from the article object. This extraction phase collects the article body, meta title and description, primary and secondary keywords, word count, internal and external link tallies, and any available SERP data such as DataForSEO results. These contextual elements form the parameter payload passed to each subsequent module in the chain.

### Stage 2: Executing the Deterministic Module Chain

The core functionality resides in the sequential execution of five analysis classes, each exposing an `analyze` method. The agent imports these from concrete implementations in the `data_sources/modules/` directory and invokes them in a strict order: **Search Intent Analyzer** → **Keyword Analyzer** → **Content Length Comparator** → **Readability Scorer** → **SEO Quality Rater**. This ordering is non-negotiable because downstream modules depend on intermediate outputs; specifically, the SEO Quality Rater requires the keyword density dictionary and link counts calculated by earlier steps.

**SearchIntentAnalyzer** ([`search_intent_analyzer.py`](https://github.com/TheCraigHewitt/seomachine/blob/main/search_intent_analyzer.py)) initiates the chain by classifying the primary keyword's intent. It uses keyword pattern matching, SERP feature mapping, and top-result content analysis to compute confidence scores. The agent calls:

```python
intent_result = SearchIntentAnalyzer().analyze(
    keyword=primary_keyword,
    serp_features=serp_features,
    top_results=top_results
)

```

**KeywordAnalyzer** ([`keyword_analyzer.py`](https://github.com/TheCraigHewitt/seomachine/blob/main/keyword_analyzer.py)) follows, calculating keyword density, distribution, and semantic clustering. Leveraging TF-IDF, K-Means clustering, and custom stop-word handling, it surfaces primary and secondary keyword usage patterns. The agent invokes:

```python
keyword_result = KeywordAnalyzer().analyze(
    content=article_content,
    primary_keyword=primary_keyword,
    secondary_keywords=secondary_keywords,
    target_density=1.5
)

```

**ContentLengthComparator** ([`content_length_comparator.py`](https://github.com/TheCraigHewitt/seomachine/blob/main/content_length_comparator.py)) benchmarks the article's word count against SERP competitors. Using `requests` and `BeautifulSoup`, it fetches top-ranking URLs, extracts word counts, and returns median and percentile statistics:

```python
length_result = ContentLengthComparator().analyze(
    keyword=primary_keyword,
    your_word_count=word_count,
    serp_results=serp_results,
    fetch_content=True
)

```

**ReadabilityScorer** ([`readability_scorer.py`](https://github.com/TheCraigHewitt/seomachine/blob/main/readability_scorer.py)) evaluates linguistic accessibility by wrapping the `textstat` library. It delivers Flesch Reading Ease scores, grade level assessments, sentence-length metrics, passive voice detection, and transition word analysis:

```python
readability_result = ReadabilityScorer().analyze(content=article_content)

```

**SeoQualityRater** ([`seo_quality_rater.py`](https://github.com/TheCraigHewitt/seomachine/blob/main/seo_quality_rater.py)) terminates the chain as the aggregation layer. It combines all previous signals with meta-element checks and link audits to emit a weighted SEO score (0-100) with category breakdowns. Crucially, it consumes the density output from KeywordAnalyzer:

```python
seo_result = SeoQualityRater().analyze(
    content=article_content,
    meta_title=meta_title,
    meta_description=meta_description,
    primary_keyword=primary_keyword,
    secondary_keywords=secondary_keywords,
    keyword_density=keyword_result['primary_keyword']['density'],
    internal_link_count=internal_links,
    external_link_count=external_links
)

```

### Stage 3: Synthesizing the Markdown Report

After all modules return Python dictionaries, the agent merges them into a richly structured Markdown report defined in the "Content Analysis Report" template within [`content-analyzer.md`](https://github.com/TheCraigHewitt/seomachine/blob/main/content-analyzer.md). This synthesis produces an executive summary, individual analytical sections for intent and readability, actionable recommendations prioritized by severity, and a competitive positioning snapshot that juxtaposes the article's metrics against SERP medians.

## Implementation: Chaining Modules in Practice

The following orchestrator script demonstrates the exact import paths and execution order used by the Content Analyzer agent. This implementation mirrors the deterministic sequence specified in the agent documentation and handles the dependency injection required by the SEO Quality Rater.

```python
from data_sources.modules.search_intent_analyzer import SearchIntentAnalyzer
from data_sources.modules.keyword_analyzer import KeywordAnalyzer
from data_sources.modules.content_length_comparator import ContentLengthComparator
from data_sources.modules.readability_scorer import ReadabilityScorer
from data_sources.modules.seo_quality_rater import SeoQualityRater

def run_content_analysis(article):
    # Extract context from article dictionary

    content = article["body"]
    meta_title = article["meta_title"]
    meta_description = article["meta_description"]
    primary_kw = article["primary_keyword"]
    secondary_kws = article.get("secondary_keywords", [])
    word_count = len(content.split())
    internal_links = article.get("internal_links", 0)
    external_links = article.get("external_links", 0)
    serp_features = article.get("serp_features")
    top_results = article.get("top_results")
    serp_results = article.get("serp_results")

    # Execute module chain in prescribed order

    intent = SearchIntentAnalyzer().analyze(
        keyword=primary_kw,
        serp_features=serp_features,
        top_results=top_results,
    )

    kw = KeywordAnalyzer().analyze(
        content=content,
        primary_keyword=primary_kw,
        secondary_keywords=secondary_kws,
        target_density=1.5,
    )

    length = ContentLengthComparator().analyze(
        keyword=primary_kw,
        your_word_count=word_count,
        serp_results=serp_results,
        fetch_content=True,
    )

    readability = ReadabilityScorer().analyze(content=content)

    # Final aggregation depends on previous outputs

    seo = SeoQualityRater().analyze(
        content=content,
        meta_title=meta_title,
        meta_description=meta_description,
        primary_keyword=primary_kw,
        secondary_keywords=secondary_kws,
        keyword_density=kw["primary_keyword"]["density"],
        internal_link_count=internal_links,
        external_link_count=external_links,
    )

    return {
        "intent": intent,
        "keyword": kw,
        "length": length,
        "readability": readability,
        "seo_score": seo,
    }

```

## Summary

- The Content Analyzer agent resides in TheCraigHewitt/seomachine repository and functions as the primary data-driven auditing component.
- Five specialized Python modules in `data_sources/modules/` perform distinct analytical tasks: intent classification, keyword optimization, length benchmarking, readability scoring, and composite SEO rating.
- The agent enforces a deterministic execution sequence—intent → keyword → length → readability → SEO rating—because the final SEO Quality Rater requires density metrics and link counts from preceding modules.
- Each module exposes a consistent `analyze` interface, enabling the agent to invoke them with specific parameter payloads extracted from the article context.
- Results synthesize into a structured Markdown report template defined in [`.claude/agents/content-analyzer.md`](https://github.com/TheCraigHewitt/seomachine/blob/main/.claude/agents/content-analyzer.md), providing actionable recommendations prioritized by severity.

## Frequently Asked Questions

### Why must the analysis modules execute in a specific order?

The **SEO Quality Rater** requires intermediate outputs generated by earlier modules in the chain. Specifically, it depends on the keyword density dictionary returned by **KeywordAnalyzer** and the internal/external link counts gathered during the initial context extraction phase. Executing modules out of sequence would result in missing dependencies and runtime errors when the final aggregation attempts to calculate the composite SEO score.

### What Python libraries power the individual analysis modules?

According to the source implementations, **KeywordAnalyzer** leverages scikit-learn's K-Means clustering and TF-IDF vectorization alongside custom stop-word handling. **ContentLengthComparator** uses `requests` for HTTP fetching and `BeautifulSoup` for HTML parsing to extract competitor word counts. **ReadabilityScorer** wraps the `textstat` library to compute Flesch Reading Ease, grade levels, and sentence complexity metrics.

### How does the agent handle missing SERP data during analysis?

The module signatures accept optional parameters such as `serp_features`, `top_results`, and `serp_results`. When these values are absent, the **SearchIntentAnalyzer** and **ContentLengthComparator** degrade gracefully, either returning intent classifications based solely on keyword pattern matching or skipping competitive length benchmarking. The agent stores these as optional dictionary keys using `.get()` methods, ensuring the pipeline continues execution without raising `KeyError` exceptions.

### Where can I modify the SEO scoring weights or add new analysis modules?

The **SeoQualityRater** implementation in [`data_sources/modules/seo_quality_rater.py`](https://github.com/TheCraigHewitt/seomachine/blob/main/data_sources/modules/seo_quality_rater.py) contains the weighted scoring logic that produces the 0-100 composite score. To adjust category weights or add new signals, modify this file's analysis method. For adding entirely new analysis dimensions, create a new module in `data_sources/modules/` exposing an `analyze` method, then insert it into the execution chain within the agent's orchestration logic before the final **SeoQualityRater** invocation.