Running the last30days-skill with Claude Code CLI: Architecture and Usage Guide
Run the last30days-skill by executing python3 scripts/last30days.py "<query>" from the repository root, which orchestrates parallel searches across Reddit, X, YouTube, and other sources, then scores and renders the results based on configurable weights for relevance, recency, and engagement.
The last30days-skill is a Claude Code skill developed by mvanhorn that automates 30-day retrospective research. It aggregates public discussions from Reddit, X (Twitter), Bluesky, Truth Social, YouTube, TikTok, Instagram, Hacker News, Polymarket, and generic web search into a ranked, de-duplicated briefing. The system runs as a single-command Python pipeline that classifies queries, executes tiered source searches in parallel, and applies a transparent scoring model to surface the most relevant content.
Installation and Environment Configuration
The skill requires API keys for external services. Configuration is loaded via scripts/lib/env.py, which checks two locations in order of precedence:
- Project-specific:
.claude/last30days.envin the repository root - Global:
~/.config/lastdays/.env
At minimum, you must provide keys for the sources you intend to query. The env.py module automatically detects available web-search backends by checking for PARALLEL_API_KEY, BRAVE_API_KEY, or OPENROUTER_API_KEY (in that priority order).
# Example .claude/last30days.env
SCRAPECREATORS_API_KEY=sk_abc123
XAI_API_KEY=xai-def456
BRAVE_API_KEY=brave-xyz789
The 8-Stage Research Pipeline
The CLI entry point in scripts/last30days.py drives an eight-stage pipeline defined in the run_research function:
- Argument Parsing: Reads global flags (e.g.,
--quick,--emit) and loads environment variables - Query Classification: Uses regex patterns in
scripts/lib/query_type.pyto categorize the query (product, concept, opinion, how-to, comparison, breaking-news, prediction) - Source Tier Selection: Determines which sources to run based on tier rules (Tier 1 always runs, Tier 2 runs if API keys exist, Tier 3 is opt-in)
- Parallel Execution: Spawns threads via
ThreadPoolExecutorto run source-specific search functions (e.g.,_search_reddit,_search_x) with timeouts enforced byTIMEOUT_PROFILES - Enrichment: Adds comment counts to Hacker News items and optionally drills into Reddit comment threads
- Entity Extraction: Phase 2 "drill-down" via
entity_extract.pypulls X handles and subreddit names from initial results to find related content missing the original keywords - Scoring & Deduplication:
scripts/lib/score.pycomputes 0-100 scores whilededupe.pyremoves duplicate URLs - Rendering:
scripts/lib/render.pyformats output ascompact(default),json,md, orcontext
Query Classification and Source Tiering
The scripts/lib/query_type.py module defines a tiered source selection system that optimizes API usage based on query intent:
- Product queries: Tier 1 runs Reddit, X, and YouTube; Tier 2 adds TikTok and Web
- Concept queries: Tier 1 runs Web and Reddit; Tier 2 adds X and YouTube
- Prediction queries: Tier 1 runs Polymarket, Reddit, and X; Tier 2 adds Web
Truth Social is always Tier 3 (opt-in via --search=truthsocial). The function is_source_enabled checks both the tier assignment and explicit user overrides via the --search flag.
The Scoring Algorithm
Each item receives a composite 0-100 score calculated in scripts/lib/score.py using weighted sub-scores:
Social Sources (Reddit/X):
WEIGHT_RELEVANCE = 0.45WEIGHT_RECENCY = 0.25WEIGHT_ENGAGEMENT = 0.30
Polymarket:
PM_WEIGHT_RELEVANCE = 0.60PM_WEIGHT_RECENCY = 0.20PM_WEIGHT_ENGAGEMENT = 0.20
Web Search:
WEBSEARCH_WEIGHT_RELEVANCE = 0.55WEBSEARCH_WEIGHT_RECENCY = 0.45- Applies
WEBSEARCH_SOURCE_PENALTY(default 15 points) unless the query type is "concept" (0 penalty)
Raw engagement values (e.g., compute_reddit_engagement_raw) normalize vote counts, comment depth, and upvote ratios to a 0-100 scale before weighting. Items missing engagement data receive an UNKNOWN_ENGAGEMENT_PENALTY of -3 points.
Running One-Shot Research Commands
Execute research directly from the shell using the main entry point:
# Basic compact output (default)
python3 scripts/last30days.py "nano banana pro prompting"
# Save as Markdown to auto-generated file path
python3 scripts/last30days.py "latest AI safety research" --emit=md
# Restrict to specific sources only
python3 scripts/last30days.py "TikTok algorithm changes" --search=reddit,x
# Use quick profile for faster results (lower timeouts, fewer results)
python3 scripts/last30days.py "iOS design trends" --quick
# Comparative analysis mode (runs parallel research passes)
python3 scripts/last30days.py "cursor vs windsurf" --emit=md
The --emit flag supports four modes: compact (one-line scored summaries), json (structured data), md (formatted briefing), and context (Claude-compatible context block).
Watchlist and Persistence Features
The open variant includes a watchlist system managed by scripts/watchlist.py and scripts/store.py:
# Add topics to periodic monitoring
python3 scripts/watchlist.py add "Claude Code updates" --frequency=weekly
python3 scripts/watchlist.py add "Polymarket AI odds" --frequency=30d
# Execute all scheduled research
python3 scripts/watchlist.py run all
When using the --store flag with the main command, results persist to ~/.config/last30days/briefings.db (SQLite), enabling historical queries like last30 what have you found about....
Core File Architecture
| File | Responsibility |
|---|---|
scripts/last30days.py |
CLI entry point and pipeline orchestration |
scripts/lib/query_type.py |
Regex classification and source tier rules |
scripts/lib/score.py |
Normalization and weighted scoring logic |
scripts/lib/env.py |
API key resolution and backend selection |
scripts/lib/entity_extract.py |
Phase 2 X handle and subreddit extraction |
scripts/lib/dedupe.py |
URL-based duplicate removal |
scripts/lib/render.py |
Output formatting for all emit modes |
scripts/lib/schema.py |
Pydantic models for type safety across sources |
scripts/lib/reddit.py |
ScrapeCreators/OpenAI-based Reddit search |
scripts/lib/bird_x.py |
X/Twitter search via Bird GraphQL or xAI |
scripts/watchlist.py |
Scheduled research management |
scripts/store.py |
SQLite persistence layer |
Summary
- The last30days-skill runs as
python3 scripts/last30days.py "<query>"and aggregates 30 days of social/web discussions - Configuration loads from
.claude/last30days.envor~/.config/last30days/.envviascripts/lib/env.py - Source selection is query-type aware, using tiers defined in
scripts/lib/query_type.pyto minimize unnecessary API calls - Scoring uses source-specific weights (e.g., 45/25/30 for Reddit/X) with penalties for web-search items to prioritize social signals
- Parallel execution with
ThreadPoolExecutorand configurableTIMEOUT_PROFILESensures responsive CLI performance - Optional watchlist functionality in
scripts/watchlist.pyenables automated periodic research with SQLite storage
Frequently Asked Questions
What API keys are required to run the last30days skill?
You only need keys for the sources you intend to query. For basic social search, provide SCRAPECREATORS_API_KEY (Reddit) and XAI_API_KEY or Bird credentials (X). For web search, provide one of PARALLEL_API_KEY, BRAVE_API_KEY, or OPENROUTER_API_KEY. The env.py module automatically detects which backends are available and skips sources with missing credentials unless explicitly requested.
How does the supplemental search phase work?
After the initial search results are collected, entity_extract.py scans content for X handles (e.g., @username) and subreddit names. The _run_supplemental function then executes focused searches for these entities in parallel. This surfaces high-engagement posts that discuss the topic without containing your original search keywords, improving recall for viral discussions and expert threads.
Why do web search results have lower scores by default?
Web search results lack native engagement metrics (likes, comments), so the scoring algorithm in scripts/lib/score.py applies a WEBSEARCH_SOURCE_PENALTY (typically 15 points) to compensate. However, the penalty varies by query type—"concept" queries receive 0 penalty because authoritative documentation is preferred, while "product" or "prediction" queries apply the full penalty to prioritize social verification and sentiment.
Can I add custom data sources to the pipeline?
Yes. Create a new module under scripts/lib/ implementing search_<source> and parse_<source>_response functions that return typed items per scripts/lib/schema.py. Then register the source in the executor block within run_research in scripts/last30days.py and add tier rules to SOURCE_TIERS in scripts/lib/query_type.py to control when it runs.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →