Implementation of TF-IDF and K-means Clustering in keyword_analyzer.py for Topic Detection
The KeywordAnalyzer class in data_sources/modules/keyword_analyzer.py implements TF-IDF vectorization and K-means clustering to automatically discover latent topics within article sections, extracting the top five terms from each cluster to generate human-readable topic labels.
The TheCraigHewitt/seomachine repository provides a comprehensive SEO analysis toolkit that includes an intelligent topic detection system. The implementation of TF-IDF and K-means clustering in keyword_analyzer.py enables the KeywordAnalyzer class to move beyond simple keyword counting to identify semantic themes within content sections. This automated topic modeling helps SEO professionals understand the thematic structure of articles and optimize internal linking strategies.
How the Topic Detection Pipeline Works
The topic detection system follows a five-step pipeline implemented in the KeywordAnalyzer class. Each step transforms raw article content into structured topic clusters that reveal the underlying thematic organization of the text.
| Step | What it does | Implementation |
|---|---|---|
| 1. Section extraction | Splits the article into meaningful sections (headings & body) so each chunk can be treated as a separate document. | _extract_sections (lines 68-115) |
| 2. TF-IDF vectorisation | Converts each section's raw text into a numerical vector that reflects how important each term is in that section and across the whole article. | TfidfVectorizer created on lines 84-88, applied on line 89 |
| 3. K-Means clustering | Groups the section vectors into a small number of clusters (max 5) where sections with similar term-distributions end up together, revealing distinct topics. | KMeans instantiated on line 95 and fitted on line 96 |
| 4. Top-terms extraction | For every cluster, the algorithm extracts the five highest-weight terms, giving a human-readable label for the discovered topic. | Loop 100-106 (uses cluster_center.argsort() and vectorizer.get_feature_names_out()) |
| 5. Result packaging | Returns a structured dictionary (clusters_found, clusters) that downstream agents (e.g., internal-linker, meta-creator) can consume. |
Return block on lines 116-119 |
Step-by-Step Implementation Details
Section Extraction with _extract_sections
Before vectorization occurs, the system must segment the article into meaningful chunks. The private method _extract_sections (lines 68-115) parses the raw markdown or HTML content to identify headings and their associated body text. Each section becomes a separate "document" for the TF-IDF vectorizer, ensuring that thematic boundaries align with the article's structural organization.
TF-IDF Vectorization Configuration
The core of the topic detection begins with scikit-learn's TfidfVectorizer, instantiated at lines 84-88:
vectorizer = TfidfVectorizer(
max_features=100,
stop_words='english',
ngram_range=(1, 2)
)
tfidf_matrix = vectorizer.fit_transform(section_texts)
Key parameters:
- max_features=100: Limits the vocabulary to the 100 most-important terms, keeping the model lightweight.
- stop_words='english': Discards common words that would otherwise dominate the vectors.
- ngram_range=(1, 2): Allows both unigrams and bigrams, so phrases like "podcast hosting" are captured as single features.
Source: lines 84-89
K-Means Clustering for Topic Grouping
Following vectorization, the implementation applies K-means clustering to group similar sections. The algorithm configuration at lines 92-96 determines the optimal number of clusters dynamically:
n_clusters = min(5, max(2, len(section_texts) // 2))
kmeans = KMeans(n_clusters=n_clusters, random_state=42, n_init=10)
cluster_labels = kmeans.fit_predict(tfidf_matrix)
Clustering logic:
- Dynamic cluster sizing: The formula
min(5, max(2, len(section_texts) // 2))ensures the analysis creates at least 2 clusters for meaningful comparison but never exceeds 5, preventing over-segmentation of short articles while allowing complex long-form content to reveal multiple distinct topics. - random_state=42: Guarantees reproducible results across multiple analysis runs.
- n_init=10: Runs the clustering algorithm 10 times with different centroid seeds and selects the best outcome, improving convergence quality.
Source: lines 92-96
Extracting Top Terms for Topic Labels
Once clusters are formed, the system extracts human-readable labels by identifying the highest-weight terms in each cluster center. The implementation at lines 100-106 sorts the cluster centers to find the dominant terms:
feature_names = vectorizer.get_feature_names_out()
for i in range(n_clusters):
cluster_center = kmeans.cluster_centers_[i]
top_indices = cluster_center.argsort()[-5:][::-1]
top_terms = [feature_names[idx] for idx in top_indices]
# ...
Label generation process:
- cluster_center: Holds the average TF-IDF weight for each term in that cluster.
- argsort()[-5:][::-1]: Identifies the indices of the five highest-weight features, reversed to show highest first.
- feature_names mapping: Converts indices back to actual words using
vectorizer.get_feature_names_out(), producing human-readable topic labels.
Source: lines 100-106
Structured Output Format
The final output (lines 116-119) packages the cluster data into a dictionary that downstream SEO agents can consume:
clusters.append({
'cluster_id': i,
'top_terms': top_terms,
'section_count': len(sections_in_cluster),
'sections': sections_in_cluster
})
The method returns a dictionary containing:
{
"clusters_found": 3,
"clusters": [
{
"cluster_id": 0,
"top_terms": ["podcast", "equipment", "microphone", "audio", "gear"],
"section_count": 3,
"sections": [0, 2, 4]
}
]
}
This structure enables internal linking modules to identify thematically related sections and meta-description generators to understand the dominant topics covered.
Source: lines 116-119
SEO Value of TF-IDF Topic Clustering
The implementation of TF-IDF and K-means clustering in keyword_analyzer.py provides several strategic advantages for SEO optimization:
- Latent topic discovery: Reveals hidden thematic structures that keyword density calculations alone cannot detect, such as distinguishing between "equipment reviews" and "hosting tutorials" within the same article.
- Content architecture validation: When sections cluster unexpectedly, it signals potential content drift or gaps in topical coverage that require editorial attention.
- Internal linking intelligence: The cluster membership data (
sectionsarray) programmatically identifies semantically related content blocks, enabling automated suggestion of internal links between topically similar sections. - Meta data optimization: The
top_termsarrays provide data-driven candidates for meta descriptions and header tags that accurately reflect the substantive topics covered.
Practical Usage Example
The following example demonstrates how to invoke the topic detection pipeline on a markdown article:
from data_sources.modules.keyword_analyzer import analyze_keywords
article = """# How to Start a Podcast
Starting a podcast is easy...
## Choosing a Topic
Pick a niche ...
## Podcast Equipment
A good microphone ...
## Hosting Platforms
Select a reliable host ...
"""
result = analyze_keywords(
content=article,
primary_keyword="start a podcast",
secondary_keywords=["podcast hosting", "podcast equipment"],
target_density=1.5,
)
# Topic clusters discovered by TF-IDF + K-Means
for cluster in result["topic_clusters"]["clusters"]:
print(f"Cluster {cluster['cluster_id']}: {', '.join(cluster['top_terms'])}")
Sample output:
Cluster 0: podcast, start, equipment, microphone, audio
Cluster 1: hosting, platform, reliable, select, choose
The clusters accurately reflect the "equipment" and "hosting" sub-topics in the article, enabling targeted SEO optimizations for each distinct theme.
Summary
- The
KeywordAnalyzerclass indata_sources/modules/keyword_analyzer.pyimplements a complete topic detection pipeline using scikit-learn's TfidfVectorizer and KMeans. - Section extraction (
_extract_sections) treats article headings as document boundaries, creating meaningful chunks for vectorization. - TF-IDF configuration uses 100 max features, English stop word removal, and 1-2 n-grams to capture both keywords and phrases.
- K-Means clustering dynamically selects 2-5 clusters based on content length, ensuring optimal topic granularity without over-segmentation.
- Top-term extraction generates human-readable labels by selecting the five highest-weight terms from each cluster center.
- The structured output enables downstream SEO agents to optimize internal linking, meta descriptions, and content architecture based on discovered topics.
Frequently Asked Questions
How does the algorithm determine the number of topics (clusters) to create?
The implementation dynamically calculates the number of clusters using the formula min(5, max(2, len(section_texts) // 2)) at line 92 in keyword_analyzer.py. This ensures the analysis creates at least 2 clusters for meaningful comparison but never exceeds 5, preventing over-segmentation of short articles while allowing complex long-form content to reveal multiple distinct topics.
What is the significance of using n-gram range (1, 2) in the TF-IDF vectorizer?
Setting ngram_range=(1, 2) on line 86 enables the vectorizer to capture both single words (unigrams) and two-word phrases (bigrams). This is critical for SEO topic detection because it preserves compound concepts like "podcast hosting" or "content strategy" as single features, preventing the semantic split that would occur if only individual words were analyzed.
How are the top terms for each cluster selected?
The algorithm extracts the top five terms by analyzing the cluster_center array at line 103, which contains the average TF-IDF weights for all terms in that cluster. It uses argsort()[-5:][::-1] to identify the indices of the five highest-weight features, then maps these indices back to actual words using vectorizer.get_feature_names_out() at line 102, producing human-readable topic labels.
Can this topic detection work without an internet connection?
Yes, the implementation is fully offline and requires no external APIs. As implemented in TheCraigHewitt/seomachine, the TF-IDF vectorization and K-means clustering rely entirely on scikit-learn's local algorithms. This makes the pipeline reproducible and scalable for batch analysis of large content libraries without incurring API costs or facing rate limits.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →