# How Query Type Detection Works in last30days-skill: A Deep Dive into the Classification Engine

> Discover how last30days-skill uses a regex-based classifier to detect query types, enabling dynamic source selection without external AI. Learn about the query type detection engine.

- Repository: [Matt Van Horn/last30days-skill](https://github.com/mvanhorn/last30days-skill)
- Tags: deep-dive
- Published: 2026-03-25

---

**The last30days-skill uses a lightweight regex-based classifier in [`scripts/lib/query_type.py`](https://github.com/mvanhorn/last30days-skill/blob/main/scripts/lib/query_type.py) to categorize every search topic into one of seven types, enabling dynamic source selection and scoring adjustments without external AI dependencies.**

The **query type detection** system sits at the heart of the **mvanhorn/last30days-skill** repository, transforming free-text user topics into structured categories that determine which data sources to query and how to rank their results. This pure-Python implementation requires no machine learning models or external API calls, making it fast, deterministic, and fully offline-capable.

## What Is Query Type Detection?

The skill defines a `Literal` type called **`QueryType`** in [`scripts/lib/query_type.py`](https://github.com/mvanhorn/last30days-skill/blob/main/scripts/lib/query_type.py) that classifies every incoming topic into one of seven distinct categories. These classifications drive downstream decisions about source selection, scoring penalties, and result ranking.

The seven supported query types are:

- **`product`** – Price-oriented or purchase-related questions (e.g., "cursor IDE pricing")
- **`concept`** – Definition and explanation requests (e.g., "what is WebTransport")
- **`opinion`** – Subjective worth and review queries (e.g., "is cursor worth it")
- **`how_to`** – Tutorial and setup instructions (e.g., "how to deploy on Vercel")
- **`comparison`** – Versus and difference queries (e.g., "cursor vs windsurf")
- **`breaking_news`** – Recent announcements and updates (e.g., "latest AI funding rounds")
- **`prediction`** – Forecasts and odds (e.g., "odds of Fed rate cut")

These values are declared at lines 6-7 of [`scripts/lib/query_type.py`](https://github.com/mvanhorn/last30days-skill/blob/main/scripts/lib/query_type.py).

## How the Detection Algorithm Works

The core function **`detect_query_type(topic: str) -> QueryType`** implements a priority-based regex matching system. At module load time, the script compiles seven case-insensitive regex patterns (lines 9-30), each designed to catch specific linguistic markers.

| Type | Regex Pattern (Case-Insensitive) | Example Triggers |
|------|----------------------------------|------------------|
| **Product** | `\b(price\|pricing\|cost\|buy\|purchase\|deal\|discount\|subscription\|plan\|tier\|free tier\|alternative\|prompt\|prompts\|prompting\|template\|templates)\b` | "best free tier LLM API" |
| **Concept** | `\b(what is\|what are\|explain\|definition\|how does\|how do\|overview\|introduction\|guide to\|primer)\b` | "explain React Server Components" |
| **Opinion** | `\b(worth it\|thoughts on\|opinion\|review\|experience with\|recommend\|should i\|pros and cons\|good or bad)\b` | "thoughts on Claude Code" |
| **How-to** | `\b(how to\|tutorial\|step by step\|setup\|install\|configure\|deploy\|migrate\|implement\|build a\|create a\|prompting\|prompts?\|best practices\|tips\|examples\|animation\|animations\|video workflow\|render pipeline)\b` | "step by step Kubernetes setup" |
| **Comparison** | `\b(vs\.?\|versus\|compared to\|comparison\|better than\|difference between\|switch from)\b` | "Claude compared to GPT-5" |
| **Breaking-news** | `\b(latest\|breaking\|just announced\|launched\|released\|new\|update\|news\|happened\|today\|this week)\b` | "OpenAI just announced GPT-6" |
| **Prediction** | `\b(predict\|forecast\|odds\|chance\|probability\|election\|outcome\|bet on\|market for)\b` | "predict the next recession" |

The detection logic follows a **priority chain** documented in the function's docstring (lines 34-38). The function checks patterns from most specific to most general, returning immediately upon the first match:

```python
if _COMPARISON_PATTERNS.search(topic):
    return "comparison"
if _HOWTO_PATTERNS.search(topic):
    return "how_to"

# ... additional checks

if _BREAKING_PATTERNS.search(topic):
    return "breaking_news"

```

If no patterns match, the function defaults to `"breaking_news"` (lines 34-38), reflecting the skill's primary use case for recent content discovery.

## Configuration Dictionaries That Drive the Pipeline

Once the type is determined, three dictionaries in [`scripts/lib/query_type.py`](https://github.com/mvanhorn/last30days-skill/blob/main/scripts/lib/query_type.py) fine-tune the retrieval pipeline:

**`SOURCE_TIERS`** (lines 62-64) maps each query type to tiered source groups. **Tier 1** sources always execute (e.g., Reddit, X, and YouTube for `product` queries). **Tier 2** sources execute only if available (e.g., general web search, TikTok). **Tier 3** sources remain opt-in only (e.g., TruthSocial).

**`WEBSEARCH_PENALTY_BY_TYPE`** (lines 75-82) adjusts scoring weights for generic web results. For example, `"concept"` queries receive a penalty of `0` because web documentation is considered authoritative for definitions, while other types may receive higher penalties to prioritize social sources.

**`TIEBREAKER_BY_TYPE`** (lines 86-94) provides deterministic ordering when multiple sources return identical scores. Lower integers indicate higher priority—for instance, a `"product"` query might prioritize Reddit with `0`, then X with `1`, then YouTube with `2`.

## Runtime Source Filtering with is_source_enabled

The helper function **`is_source_enabled(source, query_type, explicitly_requested=False)`** at lines 98-111 implements the runtime filtering logic. It consults `SOURCE_TIERS` to determine if a source belongs to Tier 1 or 2 for the given query type, and handles special opt-in rules for Tier 3 sources.

For example, TruthSocial remains disabled unless the user explicitly requests it via the `--search truthsocial` CLI flag, which passes `explicitly_requested=True` to override the tier restrictions.

## Integration Across the Codebase

The **query type detection** system propagates through multiple components according to the source analysis:

- **[`scripts/lib/reddit.py`](https://github.com/mvanhorn/last30days-skill/blob/main/scripts/lib/reddit.py)** (line 113) calls `detect_query_type` to select appropriate subreddits and adjust result scoring algorithms based on the classification.
- **[`scripts/last30days.py`](https://github.com/mvanhorn/last30days-skill/blob/main/scripts/last30days.py)** (line 1699) executes the classification immediately after parsing CLI arguments, storing the result for downstream consumers.
- **[`scripts/lib/polymarket.py`](https://github.com/mvanhorn/last30days-skill/blob/main/scripts/lib/polymarket.py)** (line 16) imports the function to adjust market data fetching strategies specifically for `prediction`-type queries.
- **[`tests/test_query_type.py`](https://github.com/mvanhorn/last30days-skill/blob/main/tests/test_query_type.py)** (lines 20-54) contains comprehensive unit tests verifying that sample topics map to expected types, ensuring regression safety for all seven categories.

## Practical Implementation Examples

Basic usage from a Python script or REPL:

```python
from scripts.lib.query_type import detect_query_type, is_source_enabled

topic = "how to deploy a Flask app on Vercel"
qtype = detect_query_type(topic)  # Returns "how_to"

print(qtype)

# Check if TikTok should run for this query type

run_tiktok = is_source_enabled("tiktok", qtype)
print(f"Run TikTok? {run_tiktok}")  # False (Tier 2 for how_to)

```

Command-line integration:

```bash
python3 scripts/last30days.py "best free tier LLM API" --emit=compact

```

Internally, the CLI executes:

```python
from scripts.lib import query_type as qt
query_type = qt.detect_query_type(args.topic)  # Determines "product"

# Source modules then check is_source_enabled(source, query_type, ...)

```

Forcing Tier 3 source opt-in:

```python

# When user explicitly requests TruthSocial via --search truthsocial

run_truth = is_source_enabled("truthsocial", qtype, explicitly_requested=True)

# Returns True regardless of default tier placement

```

## Summary

- The **query type detection** engine in [`scripts/lib/query_type.py`](https://github.com/mvanhorn/last30days-skill/blob/main/scripts/lib/query_type.py) uses compiled regex patterns to classify topics into seven categories without external dependencies.
- A **priority chain** evaluates patterns from most to least specific, falling back to `"breaking_news"` for unmatched queries.
- Three configuration dictionaries—**`SOURCE_TIERS`**, **`WEBSEARCH_PENALTY_BY_TYPE`**, and **`TIEBREAKER_BY_TYPE`**—control source selection, scoring adjustments, and tie-breaking logic.
- The **`is_source_enabled`** helper enforces tier restrictions while allowing explicit opt-in for Tier 3 sources like TruthSocial.
- Classification occurs at the CLI entry point ([`last30days.py`](https://github.com/mvanhorn/last30days-skill/blob/main/last30days.py) line 1699) and propagates to specialized fetchers like Reddit and Polymarket.

## Frequently Asked Questions

### What are the seven query types supported by last30days-skill?

The skill recognizes **product**, **concept**, **opinion**, **how_to**, **comparison**, **breaking_news**, and **prediction** types. These are defined as a `Literal` type in [`scripts/lib/query_type.py`](https://github.com/mvanhorn/last30days-skill/blob/main/scripts/lib/query_type.py) at lines 6-7, enabling static type checking while allowing string-based runtime classification.

### How does last30days-skill prioritize multiple matching query type patterns?

The `detect_query_type` function implements a **priority chain** where comparison patterns are checked first, followed by how-to, product, concept, opinion, prediction, and finally breaking-news patterns (lines 34-38). The function returns immediately upon the first match, ensuring that specific patterns like "vs" or "versus" take precedence over general news keywords.

### Can I force a specific source to run regardless of query type?

Yes. Pass `explicitly_requested=True` to the **`is_source_enabled`** function (lines 98-111). This override is used when users specify Tier 3 sources like TruthSocial via the `--search` CLI flag, forcing execution even when the source would normally be disabled for that query type.

### Where is the query type detection logic tested?

The test suite in **[`tests/test_query_type.py`](https://github.com/mvanhorn/last30days-skill/blob/main/tests/test_query_type.py)** (lines 20-54) validates the classification logic with multiple assertions covering each of the seven query types. These tests ensure that regex patterns correctly identify linguistic markers and that the priority chain resolves ambiguities consistently.