How search_intent_analyzer.py Classifies Search Queries by Intent in TheCraigHewitt/seomachine

The search_intent_analyzer.py module classifies search queries into informational, navigational, transactional, or commercial investigation intents by aggregating weighted scores from keyword pattern matching, SERP feature analysis, and top-ranking page content, then normalizing the results to determine primary and secondary intent classifications.

The search_intent_analyzer.py file in the TheCraigHewitt/seomachine repository implements a multi-source scoring system to automatically categorize search intent. By triangulating signals from the query itself, Google's SERP features, and the language of top-ranking pages, the analyzer determines whether users are seeking knowledge, trying to reach a specific website, comparing products, or ready to convert.

The Four Intent Categories

At the core of the classification system is the SearchIntent Enum defined in data_sources/modules/search_intent_analyzer.py (lines 13-19). This enumeration establishes four distinct intent categories:

  • Informational – The user seeks knowledge or answers (e.g., "how to improve SEO")
  • Navigational – The user intends to reach a specific website or page (e.g., "Facebook login")
  • Transactional – The user is ready to purchase or complete an action (e.g., "buy wireless headphones")
  • Commercial Investigation – The user is comparing options before a potential purchase (e.g., "best email marketing tools")

Signal-Based Classification Architecture

The analyzer employs a weighted scoring algorithm that evaluates three independent signal sources. Each source contributes points to one or more intent categories, and the cumulative scores determine the final classification.

Keyword Pattern Detection

The _analyze_keyword_patterns method (lines 33-67) processes the raw query text against predefined signal lists stored as class constants: INFORMATIONAL_SIGNALS, NAVIGATIONAL_SIGNALS, TRANSACTIONAL_SIGNALS, and COMMERCIAL_SIGNALS (lines 25-45).

This method applies regex-based heuristics to detect specific linguistic patterns:

  • Question words (who, what, how, why) → boosts informational scores
  • Two-word brand queries → indicates navigational intent
  • Modifiers like "best," "top," or "vs" → signals commercial investigation
  • Purchase verbs (buy, order, download) → drives transactional classification

SERP Feature Interpretation

The _analyze_serp_features method (lines 69-99) maps Google's SERP elements to intent probabilities using the SERP_INTENT_MAPPING dictionary (lines 48-58). When the analyzer receives SERP data, it iterates through feature strings and assigns weighted points based on the presence of specific elements:

  • Shopping results and ads → increase transactional confidence
  • Featured snippets and people also ask → strengthen informational classification
  • Carousels and rich results → often indicate commercial investigation

Content Pattern Analysis

The _analyze_content_patterns method (lines 100-128) examines the top 10 organic results, parsing titles, meta descriptions, and URL fragments for intent-specific vocabulary. Presence of terms like "guide," "tutorial," or "what is" boosts informational scores, while "review," "comparison," or "vs" signals commercial intent. Product-related URL paths (e.g., /buy/, /product/) or purchase-oriented descriptions increment transactional counters.

Intent Scoring and Normalization

The public analyze method (lines 61-75) serves as the entry point, accepting a keyword, optional SERP features list, and optional top-result objects. It initializes a zeroed score dictionary for all four intents, then invokes the three private scoring methods to accumulate evidence.

After aggregation, the system normalizes scores into percentage confidence values (lines 106-115). The intent with the highest raw score becomes the primary_intent, with its confidence calculated as a percentage of the total score sum. This normalization allows for direct comparison of intent strength across different queries and data richness levels.

Secondary Intent Detection

The analyzer captures nuanced search behavior by detecting secondary intents when query signals are ambiguous. If the second-highest confidence score falls within 15% of the primary intent's score (lines 117-124), that intent is recorded as secondary_intent. This threshold recognizes that queries like "best CRM software" often blend commercial investigation (comparing options) with informational (learning about features) intent.

Implementation Example

Use the analyze_intent() convenience wrapper to classify queries without instantiating the class directly:

from data_sources.modules.search_intent_analyzer import analyze_intent

# Informational query detection

info_result = analyze_intent("how to improve SEO")
print(info_result["primary_intent"])  # informational

print(info_result["confidence"])  

# {'informational': 68.0, 'commercial': 12.0, 'transactional': 10.0, 'navigational': 10.0}

# Commercial investigation with SERP data

commercial_result = analyze_intent(
    "best email marketing tools",
    serp_features=["carousel", "people_also_ask", "video"],
    top_results=[
        {"title": "Top 10 Email Marketing Tools 2024", "description": "Compare features...", "url": "https://example.com/best-email-tools"},
        {"title": "Email Marketing Software Review", "description": "Pros and cons...", "url": "https://example.com/review"}
    ]
)
print(commercial_result["primary_intent"])  # commercial

print(commercial_result["recommendations"][:2])  # Content recommendations for comparison articles

# Transactional intent with shopping signals

transactional_result = analyze_intent(
    "buy wireless headphones",
    serp_features=["shopping_results", "ads", "local_pack"]
)
print(transactional_result["primary_intent"])  # transactional

The method returns a dictionary containing the original keyword, primary_intent, optional secondary_intent, confidence percentages per category, detected signals via _get_detected_signals, and content creation recommendations via _get_recommendations.

Summary

  • The search_intent_analyzer.py module uses a multi-source scoring system combining keyword patterns, SERP features, and content analysis to classify queries.
  • Four intent types are defined in the SearchIntent Enum: informational, navigational, transactional, and commercial investigation.
  • Signal word lists and regex heuristics in _analyze_keyword_patterns detect linguistic cues within the query itself.
  • The SERP_INTENT_MAPPING dictionary translates Google's SERP elements (shopping results, carousels, snippets) into intent scores via _analyze_serp_features.
  • Top-ranking page content is parsed in _analyze_content_patterns to verify intent through titles, descriptions, and URL structures.
  • Scores are normalized to percentages, with the highest value determining the primary intent and a 15% threshold identifying secondary intents.
  • The analyze_intent() wrapper function provides a simplified functional API for integration into research pipelines.

Frequently Asked Questions

What are the four intent types defined in search_intent_analyzer.py?

The SearchIntent Enum in data_sources/modules/search_intent_analyzer.py defines informational (knowledge-seeking), navigational (site-finding), transactional (purchase-ready), and commercial investigation (comparison shopping) intents. These categories correspond to the traditional buyer's journey stages and match Google's classification of search behaviors.

How does the analyzer use SERP features to determine intent?

The _analyze_serp_features method maps SERP elements to intent scores using the SERP_INTENT_MAPPING dictionary. For example, detecting shopping_results or ads increases transactional confidence, while featured_snippet or people_also_ask signals informational intent. This approach leverages Google's own interpretation of the query, as reflected in the search results page structure.

Can search_intent_analyzer.py detect multiple intents for a single query?

Yes. The analyzer identifies a primary intent (highest normalized score) and a secondary intent when the runner-up score falls within 15% of the primary. This secondary detection captures blended intents common in modern search behavior, such as "best CRM software" which often combines commercial investigation with informational research.

Where is the intent classification logic located in the repository?

All classification logic resides in data_sources/modules/search_intent_analyzer.py. The file contains the SearchIntent Enum, signal word constants (INFORMATIONAL_SIGNALS, etc.), the SERP_INTENT_MAPPING dictionary, the core analyze method, and the analyze_intent() wrapper function. Related modules like keyword_analyzer.py consume this output for SEO metric calculations.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →