How the /analyze-existing Command Fetches, Scores, and Determines Content Health in SEOmachine

The /analyze-existing command audits existing blog posts through a three-stage pipeline: extracting content from URLs or Markdown files, processing it through five specialized analysis modules, and calculating a weighted Content-Health Score (0-100) based on six SEO pillars while flagging critical issues that block publication.

The /analyze-existing command in the TheCraigHewitt/seomachine repository is a Claude/CLI tool designed for automated content auditing. It analyzes both live web pages and local files to deliver actionable SEO recommendations and a comprehensive health assessment based on actual source code implementation.

How /analyze-existing Retrieves and Extracts Content

The command accepts two input types: a live URL or a local file path. According to the command definition in .claude/commands/analyze-existing.md (lines 18-22), the system first extracts the article's full text, headings, meta title and description, primary/secondary keywords, and publication date.

The actual fetch operation utilizes a generic fetch_content utility that handles HTML parsing for web URLs or direct file reading for local Markdown. This extraction phase prepares clean text data for the downstream analysis modules while preserving structural elements like heading hierarchy and metadata essential for SEO scoring.

The Five Analysis Modules Powering Content Scoring

Once extracted, content flows through five specialized modules located in data_sources/modules/. The Content Analyzer agent (.claude/agents/content-analyzer.md) orchestrates these components.

Search Intent Analyzer

Located at data_sources/modules/search_intent_analyzer.py, this module determines whether the content aligns with informational, transactional, or navigational search intent based on keyword context and content structure.

Keyword Analyzer

The data_sources/modules/keyword_analyzer.py module calculates keyword density, validates placement in the first 100 words, and analyzes keyword clustering to ensure optimal on-page SEO.

Content-Length Comparator

Found in data_sources/modules/content_length_comparator.py, this component benchmarks the article's word count against top SERP competitors to identify content depth gaps.

Readability Scorer

The data_sources/modules/readability_scorer.py module computes Flesch reading-ease scores, grade levels, and sentence-length metrics to ensure content accessibility matches target audience expectations.

SEO Quality Rater

The core scoring engine resides in data_sources/modules/seo_quality_rater.py. The SEOQualityRater.rate method (lines 51-84) aggregates all module outputs into the final health assessment using weighted sub-scores.

How the Content-Health Score Is Calculated

The overall_score (0-100) emerges from six weighted sub-scores defined in seo_quality_rater.py:

  • Content: 20% (0.20) — Length and depth benchmarks
  • Keywords: 25% (0.25) — Density, placement, and optimization
  • Meta: 15% (0.15) — Title and description quality
  • Structure: 15% (0.15) — Heading hierarchy and H1 presence
  • Links: 15% (0.15) — Internal and external link counts
  • Readability: 10% (0.10) — Flesch scores and sentence structure

The rate method calls private _score_* helpers to evaluate each pillar. For example, _score_structure (lines 11-16) checks for missing H1 tags, while _score_meta_elements (lines 56-60) validates meta title presence. These granular checks feed into the final weighted calculation.

What Determines Content Health?

The final health determination depends on six specific factors:

  1. Overall SEO Score — The weighted 0-100 overall_score value computed from the six pillars
  2. Category Scores — Granular metrics showing individual performance for content, keywords, meta, structure, links, and readability
  3. Critical Issues — Blockers that prevent publishing, detected in private helpers (e.g., missing H1 in _score_structure, absent keywords in first 100 words, missing meta title in _score_meta_elements)
  4. Warnings — Non-critical problems like excessive paragraph length or low internal-link counts that reduce the score
  5. Suggestions — Optimization opportunities such as adding bullet lists or increasing internal links to boost rankings
  6. Publishing-Ready Flag — In seo_quality_rater.py line 46, publishing_ready returns True only when overall_score >= 80 and zero critical issues exist

Running the Analysis Programmatically

While the /analyze-existing command orchestrates the full pipeline automatically, you can invoke the core scoring engine directly in Python:

from data_sources.modules.seo_quality_rater import rate_seo_quality
from data_sources.modules.keyword_analyzer import analyze_keywords

# Assume article_html contains fetched content

content = extract_text(article_html)  # utility that strips tags

meta_title = extract_meta_title(article_html)
meta_desc = extract_meta_description(article_html)
primary_kw = "podcast hosting"
secondary_kws = ["recording software", "microphone"]
keyword_density = 1.8  # computed via analyze_keywords

report = rate_seo_quality(
    content=content,
    meta_title=meta_title,
    meta_description=meta_desc,
    primary_keyword=primary_kw,
    secondary_keywords=secondary_kws,
    keyword_density=keyword_density,
    internal_link_count=None,   # let the rater count automatically

    external_link_count=None,
)

print(f"Overall health: {report['overall_score']}/100 – {report['grade']}")
print("Critical issues:", report["critical_issues"])
print("Quick wins:", report["suggestions"][:3])

The .claude/agents/content-analyzer.md agent normally handles this orchestration, formatting results into structured Markdown reports saved to the research/ directory.

Summary

  • The /analyze-existing command processes content through three stages: retrieval, multi-module analysis, and weighted scoring against six SEO pillars
  • Five specialized modules handle search intent, keywords, length benchmarking, readability, and SEO quality validation
  • The Content-Health Score uses specific weights: Keywords (25%), Content (20%), Meta (15%), Structure (15%), Links (15%), and Readability (10%)
  • Critical issues detected in _score_structure and _score_meta_elements helpers prevent publishing regardless of overall score
  • Content achieves publishing_ready status only when scoring 80+ with zero critical issues according to the logic in seo_quality_rater.py

Frequently Asked Questions

What input formats does the /analyze-existing command accept?

The command accepts both live URLs and local Markdown file paths. According to .claude/commands/analyze-existing.md, the extraction pipeline handles HTML parsing for web content or direct text reading for local files, extracting headings, meta data, and article text regardless of the input source.

How does the SEO Quality Rater detect critical issues?

The SEOQualityRater class uses private helper methods like _score_structure (lines 11-16) and _score_meta_elements (lines 56-60) to identify blockers. These include missing H1 tags, absent primary keywords in the first 100 words, and missing meta titles or descriptions. Each critical issue is appended to a list that prevents the publishing_ready flag from being set.

What is the minimum Content-Health Score required for publishing?

According to line 46 of data_sources/modules/seo_quality_rater.py, content receives publishing_ready = True only when the overall_score is 80 or higher and no critical issues exist in the report. Scores below 80 or any critical issue flags will mark the content as not ready for publication.

Can I customize the scoring weights for different content types?

The current implementation uses fixed weights defined in the rate method (lines 51-84) of seo_quality_rater.py: Content (20%), Keywords (25%), Meta (15%), Structure (15%), Links (15%), and Readability (10%). Modifying these values requires editing the weights dictionary within the SEOQualityRater class in the source code.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →