How Query Type Detection Works in last30days-skill: A Deep Dive into the Classification Engine
The last30days-skill uses a lightweight regex-based classifier in scripts/lib/query_type.py to categorize every search topic into one of seven types, enabling dynamic source selection and scoring adjustments without external AI dependencies.
The query type detection system sits at the heart of the mvanhorn/last30days-skill repository, transforming free-text user topics into structured categories that determine which data sources to query and how to rank their results. This pure-Python implementation requires no machine learning models or external API calls, making it fast, deterministic, and fully offline-capable.
What Is Query Type Detection?
The skill defines a Literal type called QueryType in scripts/lib/query_type.py that classifies every incoming topic into one of seven distinct categories. These classifications drive downstream decisions about source selection, scoring penalties, and result ranking.
The seven supported query types are:
product– Price-oriented or purchase-related questions (e.g., "cursor IDE pricing")concept– Definition and explanation requests (e.g., "what is WebTransport")opinion– Subjective worth and review queries (e.g., "is cursor worth it")how_to– Tutorial and setup instructions (e.g., "how to deploy on Vercel")comparison– Versus and difference queries (e.g., "cursor vs windsurf")breaking_news– Recent announcements and updates (e.g., "latest AI funding rounds")prediction– Forecasts and odds (e.g., "odds of Fed rate cut")
These values are declared at lines 6-7 of scripts/lib/query_type.py.
How the Detection Algorithm Works
The core function detect_query_type(topic: str) -> QueryType implements a priority-based regex matching system. At module load time, the script compiles seven case-insensitive regex patterns (lines 9-30), each designed to catch specific linguistic markers.
| Type | Regex Pattern (Case-Insensitive) | Example Triggers |
|---|---|---|
| Product | \b(price|pricing|cost|buy|purchase|deal|discount|subscription|plan|tier|free tier|alternative|prompt|prompts|prompting|template|templates)\b |
"best free tier LLM API" |
| Concept | \b(what is|what are|explain|definition|how does|how do|overview|introduction|guide to|primer)\b |
"explain React Server Components" |
| Opinion | \b(worth it|thoughts on|opinion|review|experience with|recommend|should i|pros and cons|good or bad)\b |
"thoughts on Claude Code" |
| How-to | \b(how to|tutorial|step by step|setup|install|configure|deploy|migrate|implement|build a|create a|prompting|prompts?|best practices|tips|examples|animation|animations|video workflow|render pipeline)\b |
"step by step Kubernetes setup" |
| Comparison | \b(vs\.?|versus|compared to|comparison|better than|difference between|switch from)\b |
"Claude compared to GPT-5" |
| Breaking-news | \b(latest|breaking|just announced|launched|released|new|update|news|happened|today|this week)\b |
"OpenAI just announced GPT-6" |
| Prediction | \b(predict|forecast|odds|chance|probability|election|outcome|bet on|market for)\b |
"predict the next recession" |
The detection logic follows a priority chain documented in the function's docstring (lines 34-38). The function checks patterns from most specific to most general, returning immediately upon the first match:
if _COMPARISON_PATTERNS.search(topic):
return "comparison"
if _HOWTO_PATTERNS.search(topic):
return "how_to"
# ... additional checks
if _BREAKING_PATTERNS.search(topic):
return "breaking_news"
If no patterns match, the function defaults to "breaking_news" (lines 34-38), reflecting the skill's primary use case for recent content discovery.
Configuration Dictionaries That Drive the Pipeline
Once the type is determined, three dictionaries in scripts/lib/query_type.py fine-tune the retrieval pipeline:
SOURCE_TIERS (lines 62-64) maps each query type to tiered source groups. Tier 1 sources always execute (e.g., Reddit, X, and YouTube for product queries). Tier 2 sources execute only if available (e.g., general web search, TikTok). Tier 3 sources remain opt-in only (e.g., TruthSocial).
WEBSEARCH_PENALTY_BY_TYPE (lines 75-82) adjusts scoring weights for generic web results. For example, "concept" queries receive a penalty of 0 because web documentation is considered authoritative for definitions, while other types may receive higher penalties to prioritize social sources.
TIEBREAKER_BY_TYPE (lines 86-94) provides deterministic ordering when multiple sources return identical scores. Lower integers indicate higher priority—for instance, a "product" query might prioritize Reddit with 0, then X with 1, then YouTube with 2.
Runtime Source Filtering with is_source_enabled
The helper function is_source_enabled(source, query_type, explicitly_requested=False) at lines 98-111 implements the runtime filtering logic. It consults SOURCE_TIERS to determine if a source belongs to Tier 1 or 2 for the given query type, and handles special opt-in rules for Tier 3 sources.
For example, TruthSocial remains disabled unless the user explicitly requests it via the --search truthsocial CLI flag, which passes explicitly_requested=True to override the tier restrictions.
Integration Across the Codebase
The query type detection system propagates through multiple components according to the source analysis:
scripts/lib/reddit.py(line 113) callsdetect_query_typeto select appropriate subreddits and adjust result scoring algorithms based on the classification.scripts/last30days.py(line 1699) executes the classification immediately after parsing CLI arguments, storing the result for downstream consumers.scripts/lib/polymarket.py(line 16) imports the function to adjust market data fetching strategies specifically forprediction-type queries.tests/test_query_type.py(lines 20-54) contains comprehensive unit tests verifying that sample topics map to expected types, ensuring regression safety for all seven categories.
Practical Implementation Examples
Basic usage from a Python script or REPL:
from scripts.lib.query_type import detect_query_type, is_source_enabled
topic = "how to deploy a Flask app on Vercel"
qtype = detect_query_type(topic) # Returns "how_to"
print(qtype)
# Check if TikTok should run for this query type
run_tiktok = is_source_enabled("tiktok", qtype)
print(f"Run TikTok? {run_tiktok}") # False (Tier 2 for how_to)
Command-line integration:
python3 scripts/last30days.py "best free tier LLM API" --emit=compact
Internally, the CLI executes:
from scripts.lib import query_type as qt
query_type = qt.detect_query_type(args.topic) # Determines "product"
# Source modules then check is_source_enabled(source, query_type, ...)
Forcing Tier 3 source opt-in:
# When user explicitly requests TruthSocial via --search truthsocial
run_truth = is_source_enabled("truthsocial", qtype, explicitly_requested=True)
# Returns True regardless of default tier placement
Summary
- The query type detection engine in
scripts/lib/query_type.pyuses compiled regex patterns to classify topics into seven categories without external dependencies. - A priority chain evaluates patterns from most to least specific, falling back to
"breaking_news"for unmatched queries. - Three configuration dictionaries—
SOURCE_TIERS,WEBSEARCH_PENALTY_BY_TYPE, andTIEBREAKER_BY_TYPE—control source selection, scoring adjustments, and tie-breaking logic. - The
is_source_enabledhelper enforces tier restrictions while allowing explicit opt-in for Tier 3 sources like TruthSocial. - Classification occurs at the CLI entry point (
last30days.pyline 1699) and propagates to specialized fetchers like Reddit and Polymarket.
Frequently Asked Questions
What are the seven query types supported by last30days-skill?
The skill recognizes product, concept, opinion, how_to, comparison, breaking_news, and prediction types. These are defined as a Literal type in scripts/lib/query_type.py at lines 6-7, enabling static type checking while allowing string-based runtime classification.
How does last30days-skill prioritize multiple matching query type patterns?
The detect_query_type function implements a priority chain where comparison patterns are checked first, followed by how-to, product, concept, opinion, prediction, and finally breaking-news patterns (lines 34-38). The function returns immediately upon the first match, ensuring that specific patterns like "vs" or "versus" take precedence over general news keywords.
Can I force a specific source to run regardless of query type?
Yes. Pass explicitly_requested=True to the is_source_enabled function (lines 98-111). This override is used when users specify Tier 3 sources like TruthSocial via the --search CLI flag, forcing execution even when the source would normally be disabled for that query type.
Where is the query type detection logic tested?
The test suite in tests/test_query_type.py (lines 20-54) validates the classification logic with multiple assertions covering each of the seven query types. These tests ensure that regex patterns correctly identify linguistic markers and that the priority chain resolves ambiguities consistently.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →