How the X vs Y Comparative Mode Works in last30days-skill

The comparative mode detects "vs" patterns using regex in scripts/lib/query_type.py, prioritizes Reddit and Hacker News over generic web results through tiered source selection and penalty scoring, and drills deeper into X handles during a second phase to surface side-by-side debates.

The last30days-skill repository by mvanhorn implements a specialized comparative query pipeline that transforms "X vs Y" questions into multi-source comparison reports. When users submit queries like "Claude vs GPT-5", the engine automatically detects the comparative intent and reconfigures its search strategy to prioritize community discussions over generic web content.

Detecting Comparative Query Patterns

The detection logic resides in scripts/lib/query_type.py where the detect_query_type() function uses a case-insensitive regex to identify comparison cues.

The pattern matches terms such as vs, versus, compared to, comparison, better than, difference between, and switch from:


# scripts/lib/query_type.py (line 23)

r"\b(vs\.?|versus|compared to|comparison|better than|difference between|switch from)\b"

When this regex matches the input string, detect_query_type() returns the literal "comparison", triggering the specialized comparative processing pipeline.

Tiered Source Selection Strategy

Once flagged as a comparison, the engine consults the SOURCE_TIERS dictionary to determine which backends to query. According to the source code in scripts/lib/query_type.py (line 67), comparative queries default to:

  • Tier 1 (always enabled): reddit, hn (Hacker News), youtube
  • Tier 2 (optional, API-dependent): x (Twitter/X), web
"comparison": {"tier1": {"reddit", "hn", "youtube"}, "tier2": {"x", "web"}},

This configuration ensures the engine always queries discussion-heavy platforms while treating generic web search and X as supplemental sources when API keys are available.

Ranking Adjustments and Penalties

The comparative mode applies specific scoring adjustments to surface community debates above standard web pages.

Web Search Penalty

The WEBSEARCH_PENALTY_BY_TYPE table in scripts/lib/query_type.py (line 80) assigns a penalty value to comparative queries:

"comparison": 10,  # mix of social and web

With a maximum penalty of 15, this score of 10 significantly reduces the weight of native web results, favoring social-media signals that typically contain richer side-by-side analyses.

Tie-Breaker Ordering

When multiple items achieve identical relevance scores, the TIEBREAKER_BY_TYPE mapping (line 92) resolves conflicts through a deterministic priority list:

"comparison": {"reddit": 0, "hn": 1, "youtube": 2, "x": 3, "web": 4, ...}

Reddit posts receive highest priority (0), followed by Hacker News (1), YouTube (2), X (3), and web results (4). This ordering reflects the comparative value of community discussions versus generic articles.

Relevance Scoring for Comparisons

The token-overlap engine in scripts/lib/relevance.py treats generic comparative terms as low-signal noise to prevent them from dominating results.

The LOW_SIGNAL_QUERY_TOKENS frozenset (lines 45-52) explicitly includes terms like 'compare', 'comparison', and 'differences':

LOW_SIGNAL_QUERY_TOKENS = frozenset({
    'advice', ..., 'compare', 'comparison', 'differences', ...
})

When a result matches only these low-signal tokens without the specific entity names (e.g., "Claude" or "GPT-5"), the relevance score is capped at 0.24, well below the typical retrieval threshold. This ensures that results explicitly mentioning both compared entities rise to the top while generic "how to compare" content is filtered out.

Phase-2 Supplemental Drilling

After the initial breadth-first search completes, the engine executes the _run_supplemental function in scripts/last30days.py to perform deep-dive retrieval.

For comparative queries, if the first pass identifies X handles (usernames) relevant to the entities being compared, the engine drills into those specific accounts:


# scripts/last30days.py (line 56-60)

has_handles = entities["x_handles"] and x_source == "bird"
...
if has_handles:
    parts.append(f"@{', @'.join(entities['x_handles'][:3])}")

This supplemental phase captures posts that discuss both entities contextually without explicitly containing the "vs" phrase, then deduplicates and merges these results into the final output.

Executing Comparative Queries

Run a comparative search from the command line:

python3 scripts/last30days.py "Claude vs GPT-5" --emit=compact

To verify query type detection programmatically:

>>> from scripts.lib import query_type as qt
>>> qt.detect_query_type("Claude compared to GPT-5")
'comparison'

Force comprehensive source coverage (all Tier 1 and Tier 2) for testing:

python3 scripts/last30days.py "Claude vs GPT-5" --sources=all --emit=json

Summary

  • Pattern Detection: scripts/lib/query_type.py uses regex to match "vs", "versus", and related terms, returning "comparison" type.
  • Source Hierarchy: Tier 1 sources (Reddit, Hacker News, YouTube) are always queried; Tier 2 (X, web) are optional.
  • Scoring Bias: Web results receive a 10/15 penalty, while Reddit gets top tie-breaker priority (0).
  • Noise Filtering: Generic terms like "compare" are capped at 0.24 relevance, ensuring entity names drive rankings.
  • Deep Retrieval: Phase-2 supplemental drilling extracts additional context from discovered X handles to enrich the comparison.

Frequently Asked Questions

How does the engine recognize an X vs Y question?

The detect_query_type() function in scripts/lib/query_type.py matches the input against a regex containing vs, versus, compared to, difference between, and similar comparative phrases. When matched, it returns the literal string "comparison", triggering the comparative mode and its associated source tiers and penalties.

Why are Reddit and Hacker News prioritized over web search for comparisons?

According to the TIEBREAKER_BY_TYPE mapping in scripts/lib/query_type.py (line 92), Reddit receives priority 0 and Hacker News priority 1, while web results receive priority 4. Additionally, the WEBSEARCH_PENALTY_BY_TYPE assigns a 10-point penalty to web results. This configuration prioritizes community platforms where side-by-side debates and user experiences are more prevalent than in generic web articles.

What happens if the query only contains generic comparison words?

The relevance engine in scripts/lib/relevance.py maintains a LOW_SIGNAL_QUERY_TOKENS set that includes "compare", "comparison", and "differences". If a result matches only these tokens without the specific entity names being compared, the relevance score is capped at 0.24. This prevents generic "how to compare" content from outranking results that explicitly mention both items in the comparison.

Can I force the comparative mode if the detector misses my query?

While the regex detection in query_type.py is automatic, you can use the --sources=all flag when running scripts/last30days.py to ensure all Tier 1 and Tier 2 sources (Reddit, Hacker News, YouTube, X, and web) are queried. This bypasses the default source selection logic and ensures comprehensive coverage regardless of the detected query type.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →