# How the /analyze-existing Command Fetches, Scores, and Determines Content Health in SEOmachine

> Learn how the SEOmachine analyze-existing command fetches, scores, and determines content health with its three-stage pipeline and Content-Health Score system.

- Repository: [Craig/seomachine](https://github.com/TheCraigHewitt/seomachine)
- Tags: how-to-guide
- Published: 2026-03-12

---

**The `/analyze-existing` command audits existing blog posts through a three-stage pipeline: extracting content from URLs or Markdown files, processing it through five specialized analysis modules, and calculating a weighted Content-Health Score (0-100) based on six SEO pillars while flagging critical issues that block publication.**

The `/analyze-existing` command in the `TheCraigHewitt/seomachine` repository is a Claude/CLI tool designed for automated content auditing. It analyzes both live web pages and local files to deliver actionable SEO recommendations and a comprehensive health assessment based on actual source code implementation.

## How /analyze-existing Retrieves and Extracts Content

The command accepts two input types: a live URL or a local file path. According to the command definition in [`.claude/commands/analyze-existing.md`](https://github.com/TheCraigHewitt/seomachine/blob/main/.claude/commands/analyze-existing.md) (lines 18-22), the system first extracts the article's **full text**, **headings**, **meta title and description**, **primary/secondary keywords**, and **publication date**. 

The actual fetch operation utilizes a generic `fetch_content` utility that handles HTML parsing for web URLs or direct file reading for local Markdown. This extraction phase prepares clean text data for the downstream analysis modules while preserving structural elements like heading hierarchy and metadata essential for SEO scoring.

## The Five Analysis Modules Powering Content Scoring

Once extracted, content flows through five specialized modules located in `data_sources/modules/`. The **Content Analyzer agent** ([`.claude/agents/content-analyzer.md`](https://github.com/TheCraigHewitt/seomachine/blob/main/.claude/agents/content-analyzer.md)) orchestrates these components.

### Search Intent Analyzer

Located at [`data_sources/modules/search_intent_analyzer.py`](https://github.com/TheCraigHewitt/seomachine/blob/main/data_sources/modules/search_intent_analyzer.py), this module determines whether the content aligns with **informational**, **transactional**, or **navigational** search intent based on keyword context and content structure.

### Keyword Analyzer

The [`data_sources/modules/keyword_analyzer.py`](https://github.com/TheCraigHewitt/seomachine/blob/main/data_sources/modules/keyword_analyzer.py) module calculates **keyword density**, validates **placement in the first 100 words**, and analyzes keyword clustering to ensure optimal on-page SEO.

### Content-Length Comparator

Found in [`data_sources/modules/content_length_comparator.py`](https://github.com/TheCraigHewitt/seomachine/blob/main/data_sources/modules/content_length_comparator.py), this component benchmarks the article's word count against top SERP competitors to identify content depth gaps.

### Readability Scorer

The [`data_sources/modules/readability_scorer.py`](https://github.com/TheCraigHewitt/seomachine/blob/main/data_sources/modules/readability_scorer.py) module computes **Flesch reading-ease scores**, grade levels, and sentence-length metrics to ensure content accessibility matches target audience expectations.

### SEO Quality Rater

The core scoring engine resides in [`data_sources/modules/seo_quality_rater.py`](https://github.com/TheCraigHewitt/seomachine/blob/main/data_sources/modules/seo_quality_rater.py). The `SEOQualityRater.rate` method (lines 51-84) aggregates all module outputs into the final health assessment using weighted sub-scores.

## How the Content-Health Score Is Calculated

The `overall_score` (0-100) emerges from six weighted sub-scores defined in [`seo_quality_rater.py`](https://github.com/TheCraigHewitt/seomachine/blob/main/seo_quality_rater.py):

- **Content**: 20% (0.20) — Length and depth benchmarks
- **Keywords**: 25% (0.25) — Density, placement, and optimization
- **Meta**: 15% (0.15) — Title and description quality
- **Structure**: 15% (0.15) — Heading hierarchy and H1 presence
- **Links**: 15% (0.15) — Internal and external link counts
- **Readability**: 10% (0.10) — Flesch scores and sentence structure

The `rate` method calls private `_score_*` helpers to evaluate each pillar. For example, `_score_structure` (lines 11-16) checks for missing H1 tags, while `_score_meta_elements` (lines 56-60) validates meta title presence. These granular checks feed into the final weighted calculation.

## What Determines Content Health?

The final health determination depends on six specific factors:

1. **Overall SEO Score** — The weighted 0-100 `overall_score` value computed from the six pillars
2. **Category Scores** — Granular metrics showing individual performance for content, keywords, meta, structure, links, and readability
3. **Critical Issues** — Blockers that prevent publishing, detected in private helpers (e.g., missing H1 in `_score_structure`, absent keywords in first 100 words, missing meta title in `_score_meta_elements`)
4. **Warnings** — Non-critical problems like excessive paragraph length or low internal-link counts that reduce the score
5. **Suggestions** — Optimization opportunities such as adding bullet lists or increasing internal links to boost rankings
6. **Publishing-Ready Flag** — In [`seo_quality_rater.py`](https://github.com/TheCraigHewitt/seomachine/blob/main/seo_quality_rater.py) line 46, `publishing_ready` returns `True` only when `overall_score >= 80` **and** zero critical issues exist

## Running the Analysis Programmatically

While the `/analyze-existing` command orchestrates the full pipeline automatically, you can invoke the core scoring engine directly in Python:

```python
from data_sources.modules.seo_quality_rater import rate_seo_quality
from data_sources.modules.keyword_analyzer import analyze_keywords

# Assume article_html contains fetched content

content = extract_text(article_html)  # utility that strips tags

meta_title = extract_meta_title(article_html)
meta_desc = extract_meta_description(article_html)
primary_kw = "podcast hosting"
secondary_kws = ["recording software", "microphone"]
keyword_density = 1.8  # computed via analyze_keywords

report = rate_seo_quality(
    content=content,
    meta_title=meta_title,
    meta_description=meta_desc,
    primary_keyword=primary_kw,
    secondary_keywords=secondary_kws,
    keyword_density=keyword_density,
    internal_link_count=None,   # let the rater count automatically

    external_link_count=None,
)

print(f"Overall health: {report['overall_score']}/100 – {report['grade']}")
print("Critical issues:", report["critical_issues"])
print("Quick wins:", report["suggestions"][:3])

```

The [`.claude/agents/content-analyzer.md`](https://github.com/TheCraigHewitt/seomachine/blob/main/.claude/agents/content-analyzer.md) agent normally handles this orchestration, formatting results into structured Markdown reports saved to the `research/` directory.

## Summary

- The `/analyze-existing` command processes content through three stages: retrieval, multi-module analysis, and weighted scoring against six SEO pillars
- Five specialized modules handle **search intent**, **keywords**, **length benchmarking**, **readability**, and **SEO quality** validation
- The Content-Health Score uses specific weights: **Keywords (25%)**, **Content (20%)**, **Meta (15%)**, **Structure (15%)**, **Links (15%)**, and **Readability (10%)**
- Critical issues detected in `_score_structure` and `_score_meta_elements` helpers prevent publishing regardless of overall score
- Content achieves `publishing_ready` status only when scoring **80+** with **zero critical issues** according to the logic in [`seo_quality_rater.py`](https://github.com/TheCraigHewitt/seomachine/blob/main/seo_quality_rater.py)

## Frequently Asked Questions

### What input formats does the /analyze-existing command accept?

The command accepts both live URLs and local Markdown file paths. According to [`.claude/commands/analyze-existing.md`](https://github.com/TheCraigHewitt/seomachine/blob/main/.claude/commands/analyze-existing.md), the extraction pipeline handles HTML parsing for web content or direct text reading for local files, extracting headings, meta data, and article text regardless of the input source.

### How does the SEO Quality Rater detect critical issues?

The `SEOQualityRater` class uses private helper methods like `_score_structure` (lines 11-16) and `_score_meta_elements` (lines 56-60) to identify blockers. These include missing H1 tags, absent primary keywords in the first 100 words, and missing meta titles or descriptions. Each critical issue is appended to a list that prevents the `publishing_ready` flag from being set.

### What is the minimum Content-Health Score required for publishing?

According to line 46 of [`data_sources/modules/seo_quality_rater.py`](https://github.com/TheCraigHewitt/seomachine/blob/main/data_sources/modules/seo_quality_rater.py), content receives `publishing_ready = True` only when the `overall_score` is **80 or higher** and **no critical issues** exist in the report. Scores below 80 or any critical issue flags will mark the content as not ready for publication.

### Can I customize the scoring weights for different content types?

The current implementation uses fixed weights defined in the `rate` method (lines 51-84) of [`seo_quality_rater.py`](https://github.com/TheCraigHewitt/seomachine/blob/main/seo_quality_rater.py): Content (20%), Keywords (25%), Meta (15%), Structure (15%), Links (15%), and Readability (10%). Modifying these values requires editing the weights dictionary within the `SEOQualityRater` class in the source code.