How the Scoring and Ranking Algorithm Works in last30days-skill

The last30days-skill calculates a 0-100 quality score for every search result by combining relevance, recency, and engagement signals with source-specific weights, then ranks items by score descending while applying penalties for missing data and low date confidence.

The last30days-skill (available in the mvanhorn/last30days-skill repository) aggregates content from Reddit, X, YouTube, TikTok, and other platforms to surface the most pertinent information from the past 30 days. Its scoring and ranking algorithm evaluates each item through a multi-dimensional lens, blending AI-assessed relevance with temporal freshness and community engagement metrics before producing a final sorted list for the user.

The Three Sub-Scores

Every item receives three normalized sub-scores (0-100) that capture distinct quality signals.

Relevance Score

The relevance sub-score reflects how well the content matches the user's query. The system extracts this directly from the source's relevance field (typically a 0-1 float provided by the search model) and scales it to an integer percentage:

relevance_score = int(item.relevance * 100)

Recency Score

The recency sub-score measures freshness within the 30-day window. The algorithm converts date strings into "days ago" values and applies a linear decay function. Items published today receive 100, while items at the 30-day boundary approach 0. This calculation lives in scripts/lib/dates.py within the recency_score function.

Engagement Score

The engagement sub-score estimates popularity using platform-specific interaction metrics. Raw signals (views, likes, comments, upvotes) are first transformed with log1p to dampen outliers, then blended with source-specific coefficients. The final raw engagement value is normalized to a 0-100 scale across all items in the result set using normalize_to_100 (see scripts/lib/score.py lines 87-117).

Source-Specific Weighting Formulas

The algorithm combines the three sub-scores using weights that vary by source type to account for platform-specific signal strengths.

Default Social Platform Weights

For Reddit, X, YouTube, and most social sources, the default weighting in scripts/lib/score.py (lines 9-13) is:

WEIGHT_RELEVANCE = 0.45   # 45%

WEIGHT_RECENCY   = 0.25   # 25%

WEIGHT_ENGAGEMENT = 0.30  # 30%

Polymarket Weights

Because market volume already signals interest, Polymarket uses higher relevance weighting (lines 14-18):

PM_WEIGHT_RELEVANCE = 0.60
PM_WEIGHT_RECENCY   = 0.20
PM_WEIGHT_ENGAGEMENT = 0.20

Platform-Specific Engagement Blends

Each source defines its own raw engagement formula in scripts/lib/score.py using log1p to safely handle zero or missing values:

  • Reddit (lines 43-66): 0.50*log1p(score) + 0.35*log1p(num_comments) + 0.05*(upvote_ratio*10) + 0.10*log1p(top_comment_score)
  • X (lines 68-84): 0.55*log1p(likes) + 0.25*log1p(reposts) + 0.15*log1p(replies) + 0.05*log1p(quotes)
  • YouTube (lines 45-61): 0.50*log1p(views) + 0.35*log1p(likes) + 0.15*log1p(comments)

Normalization and Score Adjustments

After calculating the weighted sum, the algorithm applies several penalties to refine accuracy.

Missing Engagement Penalty

Items lacking engagement data receive a flat deduction defined by UNKNOWN_ENGAGEMENT_PENALTY = 3 points (applied after weighting, see lines 31-34).

Date Confidence Penalty

The date_confidence field (high/med/low) adjusts scores to reflect temporal uncertainty. Low confidence reduces the score by 5 points; medium confidence reduces it by 2 points (lines 74-78).

WebSearch Type Penalty

Generic web search results incur a WEBSEARCH_PENALTY_BY_TYPE adjustment (default 15 points) applied before the final weighting to prioritize curated social content over raw web pages (lines 72-83).

Final Score Clamping

The algorithm ensures all scores fall within bounds:

item.score = max(0, min(100, int(overall)))

This logic appears in each source-specific scoring function, such as score_reddit_items (lines 164-180).

Final Ranking and Sorting Logic

Once individual items are scored, the system constructs a unified ranking across all sources.

Flat List Construction

The build_ranked_items function in scripts/evaluate_search_quality.py (lines 14-29) collects the top N items per source (controlled by per_source_limit) into a single flat list.

Multi-Level Sorting

The final sort uses a three-tier tuple key for deterministic ordering:

  1. Score descending (primary sort)
  2. Source name (secondary tie-breaker)
  3. Stable item key (tertiary tie-breaker)
ranked.sort(key=lambda item: (-item["score"], item["source"], item["key"]))

See lines 28-29 in scripts/evaluate_search_quality.py for the implementation.

Query-Type Awareness

The skill adjusts scoring parameters based on query intent detected via detect_query_type in scripts/lib/query_type.py (lines 8-30). The system classifies queries into seven types: product, concept, opinion, how_to, comparison, breaking_news, and prediction.

Impact on Results

Query type influences three scoring dimensions:

  • Source tiering: Determines which sources execute via SOURCE_TIERS (lines 62-70)
  • WebSearch penalty: Reduces severity for "concept" or "how_to" queries (lines 72-82)
  • Tie-breaker priority: Adjusts deterministic ordering via TIEBREAKER_BY_TYPE (lines 85-95)

Implementation Examples

Scoring Reddit Items

To score a batch of Reddit posts programmatically:

from scripts.lib import score, schema

# Assume items is a list of schema.RedditItem objects

scored_items = score.score_reddit_items(items)

for item in scored_items:
    print(item.title, item.score, item.subs)

The subs attribute contains the raw sub-scores for relevance, recency, and engagement.

Building a Ranked Result Set

To merge and rank results from multiple sources:

from scripts.evaluate_search_quality import build_ranked_items

# report is the dict produced by the search engine, keyed by source

ranked = build_ranked_items(report, per_source_limit=5)

for entry in ranked:
    print(entry["source"], entry["score"], entry["text"])

Summary

  • The scoring pipeline in scripts/lib/score.py calculates three sub-scores (relevance, recency, engagement) and combines them with source-specific weights (45/25/30 for most platforms, 60/20/20 for Polymarket).
  • Raw engagement uses log1p transformations with platform-specific coefficients to prevent outliers from dominating results.
  • Final scores are adjusted with penalties for missing engagement (-3), low date confidence (-5), and generic web content (-15) before clamping to 0-100.
  • The ranking step in scripts/evaluate_search_quality.py sorts by descending score, then source name, then stable key for deterministic tie-breaking.
  • Query-type detection dynamically adjusts source selection, penalty severity, and tie-breaker logic to match user intent.

Frequently Asked Questions

How does last30days-skill handle items with missing engagement data?

Items missing engagement metrics receive a penalty of 3 points after the initial weighting calculation. The log1p function used in raw engagement formulas ensures that zero or null values become 0 rather than causing errors, though they will significantly lower the engagement sub-score component.

Why does Polymarket use different scoring weights than Reddit or X?

Polymarket uses 60% relevance weighting versus the standard 45% because market volume already serves as a strong proxy for engagement. This configuration, defined in scripts/lib/score.py lines 14-18, allows the algorithm to prioritize semantic relevance while still considering market activity and recency.

What determines the final sort order when two items have identical scores?

When scores are identical, the algorithm applies a deterministic secondary sort by source name, followed by a tertiary sort on a stable item identifier. This three-tier sorting key ensures consistent, reproducible rankings across identical queries, as implemented in build_ranked_items at lines 28-29 of scripts/evaluate_search_quality.py.

How does the date confidence field affect an item's final score?

The date_confidence field (high, medium, or low) applies a post-weighting penalty: low confidence reduces the score by 5 points, while medium confidence reduces it by 2 points. This adjustment compensates for temporal uncertainty when parsing publication dates from unstructured sources.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →