# Implementation of TF-IDF and K-means Clustering in keyword_analyzer.py for Topic Detection

> Explore the implementation of TF-IDF and K-means clustering in keyword_analyzer.py for automatic topic detection. Discover latent topics and generate readable labels.

- Repository: [Craig/seomachine](https://github.com/TheCraigHewitt/seomachine)
- Tags: how-to-guide
- Published: 2026-03-12

---

**The `KeywordAnalyzer` class in [`data_sources/modules/keyword_analyzer.py`](https://github.com/TheCraigHewitt/seomachine/blob/main/data_sources/modules/keyword_analyzer.py) implements TF-IDF vectorization and K-means clustering to automatically discover latent topics within article sections, extracting the top five terms from each cluster to generate human-readable topic labels.**

The TheCraigHewitt/seomachine repository provides a comprehensive SEO analysis toolkit that includes an intelligent topic detection system. The implementation of TF-IDF and K-means clustering in [`keyword_analyzer.py`](https://github.com/TheCraigHewitt/seomachine/blob/main/keyword_analyzer.py) enables the `KeywordAnalyzer` class to move beyond simple keyword counting to identify semantic themes within content sections. This automated topic modeling helps SEO professionals understand the thematic structure of articles and optimize internal linking strategies.

## How the Topic Detection Pipeline Works

The topic detection system follows a five-step pipeline implemented in the `KeywordAnalyzer` class. Each step transforms raw article content into structured topic clusters that reveal the underlying thematic organization of the text.

| Step | What it does | Implementation |
|------|--------------|----------------|
| **1. Section extraction** | Splits the article into meaningful sections (headings & body) so each chunk can be treated as a separate document. | `_extract_sections` (lines 68-115) |
| **2. TF-IDF vectorisation** | Converts each section's raw text into a numerical vector that reflects how important each term is in that section *and* across the whole article. | `TfidfVectorizer` created on lines 84-88, applied on line 89 |
| **3. K-Means clustering** | Groups the section vectors into a small number of clusters (max 5) where sections with similar term-distributions end up together, revealing distinct topics. | `KMeans` instantiated on line 95 and fitted on line 96 |
| **4. Top-terms extraction** | For every cluster, the algorithm extracts the five highest-weight terms, giving a human-readable label for the discovered topic. | Loop 100-106 (uses `cluster_center.argsort()` and `vectorizer.get_feature_names_out()`) |
| **5. Result packaging** | Returns a structured dictionary (`clusters_found`, `clusters`) that downstream agents (e.g., internal-linker, meta-creator) can consume. | Return block on lines 116-119 |

## Step-by-Step Implementation Details

### Section Extraction with `_extract_sections`

Before vectorization occurs, the system must segment the article into meaningful chunks. The private method `_extract_sections` (lines 68-115) parses the raw markdown or HTML content to identify headings and their associated body text. Each section becomes a separate "document" for the TF-IDF vectorizer, ensuring that thematic boundaries align with the article's structural organization.

### TF-IDF Vectorization Configuration

The core of the topic detection begins with scikit-learn's `TfidfVectorizer`, instantiated at lines 84-88:

```python
vectorizer = TfidfVectorizer(
    max_features=100,
    stop_words='english',
    ngram_range=(1, 2)
)
tfidf_matrix = vectorizer.fit_transform(section_texts)

```

**Key parameters:**

- **max_features=100**: Limits the vocabulary to the 100 most-important terms, keeping the model lightweight.
- **stop_words='english'**: Discards common words that would otherwise dominate the vectors.
- **ngram_range=(1, 2)**: Allows both unigrams and bigrams, so phrases like "podcast hosting" are captured as single features.

Source: [lines 84-89](https://github.com/TheCraigHewitt/seomachine/blob/main/data_sources/modules/keyword_analyzer.py#L84-L89)

### K-Means Clustering for Topic Grouping

Following vectorization, the implementation applies K-means clustering to group similar sections. The algorithm configuration at lines 92-96 determines the optimal number of clusters dynamically:

```python
n_clusters = min(5, max(2, len(section_texts) // 2))
kmeans = KMeans(n_clusters=n_clusters, random_state=42, n_init=10)
cluster_labels = kmeans.fit_predict(tfidf_matrix)

```

**Clustering logic:**

- **Dynamic cluster sizing**: The formula `min(5, max(2, len(section_texts) // 2))` ensures the analysis creates at least 2 clusters for meaningful comparison but never exceeds 5, preventing over-segmentation of short articles while allowing complex long-form content to reveal multiple distinct topics.
- **random_state=42**: Guarantees reproducible results across multiple analysis runs.
- **n_init=10**: Runs the clustering algorithm 10 times with different centroid seeds and selects the best outcome, improving convergence quality.

Source: [lines 92-96](https://github.com/TheCraigHewitt/seomachine/blob/main/data_sources/modules/keyword_analyzer.py#L92-L96)

### Extracting Top Terms for Topic Labels

Once clusters are formed, the system extracts human-readable labels by identifying the highest-weight terms in each cluster center. The implementation at lines 100-106 sorts the cluster centers to find the dominant terms:

```python
feature_names = vectorizer.get_feature_names_out()
for i in range(n_clusters):
    cluster_center = kmeans.cluster_centers_[i]
    top_indices = cluster_center.argsort()[-5:][::-1]
    top_terms = [feature_names[idx] for idx in top_indices]
    # ...

```

**Label generation process:**

- **cluster_center**: Holds the average TF-IDF weight for each term in that cluster.
- **argsort()[-5:][::-1]**: Identifies the indices of the five highest-weight features, reversed to show highest first.
- **feature_names mapping**: Converts indices back to actual words using `vectorizer.get_feature_names_out()`, producing human-readable topic labels.

Source: [lines 100-106](https://github.com/TheCraigHewitt/seomachine/blob/main/data_sources/modules/keyword_analyzer.py#L100-L106)

### Structured Output Format

The final output (lines 116-119) packages the cluster data into a dictionary that downstream SEO agents can consume:

```python
clusters.append({
    'cluster_id': i,
    'top_terms': top_terms,
    'section_count': len(sections_in_cluster),
    'sections': sections_in_cluster
})

```

The method returns a dictionary containing:

```json
{
  "clusters_found": 3,
  "clusters": [
    {
      "cluster_id": 0,
      "top_terms": ["podcast", "equipment", "microphone", "audio", "gear"],
      "section_count": 3,
      "sections": [0, 2, 4]
    }
  ]
}

```

This structure enables internal linking modules to identify thematically related sections and meta-description generators to understand the dominant topics covered.

Source: [lines 116-119](https://github.com/TheCraigHewitt/seomachine/blob/main/data_sources/modules/keyword_analyzer.py#L116-L119)

## SEO Value of TF-IDF Topic Clustering

The implementation of TF-IDF and K-means clustering in [`keyword_analyzer.py`](https://github.com/TheCraigHewitt/seomachine/blob/main/keyword_analyzer.py) provides several strategic advantages for SEO optimization:

- **Latent topic discovery**: Reveals hidden thematic structures that keyword density calculations alone cannot detect, such as distinguishing between "equipment reviews" and "hosting tutorials" within the same article.
- **Content architecture validation**: When sections cluster unexpectedly, it signals potential content drift or gaps in topical coverage that require editorial attention.
- **Internal linking intelligence**: The cluster membership data (`sections` array) programmatically identifies semantically related content blocks, enabling automated suggestion of internal links between topically similar sections.
- **Meta data optimization**: The `top_terms` arrays provide data-driven candidates for meta descriptions and header tags that accurately reflect the substantive topics covered.

## Practical Usage Example

The following example demonstrates how to invoke the topic detection pipeline on a markdown article:

```python
from data_sources.modules.keyword_analyzer import analyze_keywords

article = """# How to Start a Podcast

Starting a podcast is easy...

## Choosing a Topic

Pick a niche ...

## Podcast Equipment

A good microphone ...

## Hosting Platforms

Select a reliable host ...
"""

result = analyze_keywords(
    content=article,
    primary_keyword="start a podcast",
    secondary_keywords=["podcast hosting", "podcast equipment"],
    target_density=1.5,
)

# Topic clusters discovered by TF-IDF + K-Means

for cluster in result["topic_clusters"]["clusters"]:
    print(f"Cluster {cluster['cluster_id']}: {', '.join(cluster['top_terms'])}")

```

**Sample output:**

```

Cluster 0: podcast, start, equipment, microphone, audio
Cluster 1: hosting, platform, reliable, select, choose

```

The clusters accurately reflect the "equipment" and "hosting" sub-topics in the article, enabling targeted SEO optimizations for each distinct theme.

## Summary

- The `KeywordAnalyzer` class in [`data_sources/modules/keyword_analyzer.py`](https://github.com/TheCraigHewitt/seomachine/blob/main/data_sources/modules/keyword_analyzer.py) implements a complete topic detection pipeline using scikit-learn's **TfidfVectorizer** and **KMeans**.
- **Section extraction** (`_extract_sections`) treats article headings as document boundaries, creating meaningful chunks for vectorization.
- **TF-IDF configuration** uses 100 max features, English stop word removal, and 1-2 n-grams to capture both keywords and phrases.
- **K-Means clustering** dynamically selects 2-5 clusters based on content length, ensuring optimal topic granularity without over-segmentation.
- **Top-term extraction** generates human-readable labels by selecting the five highest-weight terms from each cluster center.
- The structured output enables downstream SEO agents to optimize internal linking, meta descriptions, and content architecture based on discovered topics.

## Frequently Asked Questions

### How does the algorithm determine the number of topics (clusters) to create?

The implementation dynamically calculates the number of clusters using the formula `min(5, max(2, len(section_texts) // 2))` at line 92 in [`keyword_analyzer.py`](https://github.com/TheCraigHewitt/seomachine/blob/main/keyword_analyzer.py). This ensures the analysis creates at least 2 clusters for meaningful comparison but never exceeds 5, preventing over-segmentation of short articles while allowing complex long-form content to reveal multiple distinct topics.

### What is the significance of using n-gram range (1, 2) in the TF-IDF vectorizer?

Setting `ngram_range=(1, 2)` on line 86 enables the vectorizer to capture both single words (unigrams) and two-word phrases (bigrams). This is critical for SEO topic detection because it preserves compound concepts like "podcast hosting" or "content strategy" as single features, preventing the semantic split that would occur if only individual words were analyzed.

### How are the top terms for each cluster selected?

The algorithm extracts the top five terms by analyzing the `cluster_center` array at line 103, which contains the average TF-IDF weights for all terms in that cluster. It uses `argsort()[-5:][::-1]` to identify the indices of the five highest-weight features, then maps these indices back to actual words using `vectorizer.get_feature_names_out()` at line 102, producing human-readable topic labels.

### Can this topic detection work without an internet connection?

Yes, the implementation is fully offline and requires no external APIs. As implemented in TheCraigHewitt/seomachine, the TF-IDF vectorization and K-means clustering rely entirely on scikit-learn's local algorithms. This makes the pipeline reproducible and scalable for batch analysis of large content libraries without incurring API costs or facing rate limits.