# Running the last30days-skill with Claude Code CLI: Architecture and Usage Guide

> Easily run the last30days-skill with Codex CLI. Orchestrate parallel searches across multiple sources, score results by relevance & recency, and render them with this guide.

- Repository: [Matt Van Horn/last30days-skill](https://github.com/mvanhorn/last30days-skill)
- Tags: architecture
- Published: 2026-03-25

---

**Run the last30days-skill by executing `python3 scripts/last30days.py "<query>"` from the repository root, which orchestrates parallel searches across Reddit, X, YouTube, and other sources, then scores and renders the results based on configurable weights for relevance, recency, and engagement.**

The `last30days-skill` is a **Claude Code skill** developed by `mvanhorn` that automates 30-day retrospective research. It aggregates public discussions from Reddit, X (Twitter), Bluesky, Truth Social, YouTube, TikTok, Instagram, Hacker News, Polymarket, and generic web search into a ranked, de-duplicated briefing. The system runs as a single-command Python pipeline that classifies queries, executes tiered source searches in parallel, and applies a transparent scoring model to surface the most relevant content.

## Installation and Environment Configuration

The skill requires API keys for external services. Configuration is loaded via [`scripts/lib/env.py`](https://github.com/mvanhorn/last30days-skill/blob/main/scripts/lib/env.py), which checks two locations in order of precedence:

1. **Project-specific**: `.claude/last30days.env` in the repository root
2. **Global**: `~/.config/lastdays/.env`

At minimum, you must provide keys for the sources you intend to query. The [`env.py`](https://github.com/mvanhorn/last30days-skill/blob/main/env.py) module automatically detects available web-search backends by checking for `PARALLEL_API_KEY`, `BRAVE_API_KEY`, or `OPENROUTER_API_KEY` (in that priority order).

```bash

# Example .claude/last30days.env

SCRAPECREATORS_API_KEY=sk_abc123
XAI_API_KEY=xai-def456
BRAVE_API_KEY=brave-xyz789

```

## The 8-Stage Research Pipeline

The CLI entry point in [`scripts/last30days.py`](https://github.com/mvanhorn/last30days-skill/blob/main/scripts/last30days.py) drives an eight-stage pipeline defined in the `run_research` function:

1. **Argument Parsing**: Reads global flags (e.g., `--quick`, `--emit`) and loads environment variables
2. **Query Classification**: Uses regex patterns in [`scripts/lib/query_type.py`](https://github.com/mvanhorn/last30days-skill/blob/main/scripts/lib/query_type.py) to categorize the query (product, concept, opinion, how-to, comparison, breaking-news, prediction)
3. **Source Tier Selection**: Determines which sources to run based on tier rules (Tier 1 always runs, Tier 2 runs if API keys exist, Tier 3 is opt-in)
4. **Parallel Execution**: Spawns threads via `ThreadPoolExecutor` to run source-specific search functions (e.g., `_search_reddit`, `_search_x`) with timeouts enforced by `TIMEOUT_PROFILES`
5. **Enrichment**: Adds comment counts to Hacker News items and optionally drills into Reddit comment threads
6. **Entity Extraction**: Phase 2 "drill-down" via [`entity_extract.py`](https://github.com/mvanhorn/last30days-skill/blob/main/entity_extract.py) pulls X handles and subreddit names from initial results to find related content missing the original keywords
7. **Scoring & Deduplication**: [`scripts/lib/score.py`](https://github.com/mvanhorn/last30days-skill/blob/main/scripts/lib/score.py) computes 0-100 scores while [`dedupe.py`](https://github.com/mvanhorn/last30days-skill/blob/main/dedupe.py) removes duplicate URLs
8. **Rendering**: [`scripts/lib/render.py`](https://github.com/mvanhorn/last30days-skill/blob/main/scripts/lib/render.py) formats output as `compact` (default), `json`, `md`, or `context`

## Query Classification and Source Tiering

The [`scripts/lib/query_type.py`](https://github.com/mvanhorn/last30days-skill/blob/main/scripts/lib/query_type.py) module defines a tiered source selection system that optimizes API usage based on query intent:

- **Product queries**: Tier 1 runs Reddit, X, and YouTube; Tier 2 adds TikTok and Web
- **Concept queries**: Tier 1 runs Web and Reddit; Tier 2 adds X and YouTube
- **Prediction queries**: Tier 1 runs Polymarket, Reddit, and X; Tier 2 adds Web

Truth Social is always Tier 3 (opt-in via `--search=truthsocial`). The function `is_source_enabled` checks both the tier assignment and explicit user overrides via the `--search` flag.

## The Scoring Algorithm

Each item receives a composite 0-100 score calculated in [`scripts/lib/score.py`](https://github.com/mvanhorn/last30days-skill/blob/main/scripts/lib/score.py) using weighted sub-scores:

**Social Sources (Reddit/X)**:
- `WEIGHT_RELEVANCE = 0.45`
- `WEIGHT_RECENCY = 0.25`  
- `WEIGHT_ENGAGEMENT = 0.30`

**Polymarket**:
- `PM_WEIGHT_RELEVANCE = 0.60`
- `PM_WEIGHT_RECENCY = 0.20`
- `PM_WEIGHT_ENGAGEMENT = 0.20`

**Web Search**:
- `WEBSEARCH_WEIGHT_RELEVANCE = 0.55`
- `WEBSEARCH_WEIGHT_RECENCY = 0.45`
- Applies `WEBSEARCH_SOURCE_PENALTY` (default 15 points) unless the query type is "concept" (0 penalty)

Raw engagement values (e.g., `compute_reddit_engagement_raw`) normalize vote counts, comment depth, and upvote ratios to a 0-100 scale before weighting. Items missing engagement data receive an `UNKNOWN_ENGAGEMENT_PENALTY` of -3 points.

## Running One-Shot Research Commands

Execute research directly from the shell using the main entry point:

```bash

# Basic compact output (default)

python3 scripts/last30days.py "nano banana pro prompting"

# Save as Markdown to auto-generated file path

python3 scripts/last30days.py "latest AI safety research" --emit=md

# Restrict to specific sources only

python3 scripts/last30days.py "TikTok algorithm changes" --search=reddit,x

# Use quick profile for faster results (lower timeouts, fewer results)

python3 scripts/last30days.py "iOS design trends" --quick

# Comparative analysis mode (runs parallel research passes)

python3 scripts/last30days.py "cursor vs windsurf" --emit=md

```

The `--emit` flag supports four modes: `compact` (one-line scored summaries), `json` (structured data), `md` (formatted briefing), and `context` (Claude-compatible context block).

## Watchlist and Persistence Features

The open variant includes a watchlist system managed by [`scripts/watchlist.py`](https://github.com/mvanhorn/last30days-skill/blob/main/scripts/watchlist.py) and [`scripts/store.py`](https://github.com/mvanhorn/last30days-skill/blob/main/scripts/store.py):

```bash

# Add topics to periodic monitoring

python3 scripts/watchlist.py add "Claude Code updates" --frequency=weekly
python3 scripts/watchlist.py add "Polymarket AI odds" --frequency=30d

# Execute all scheduled research

python3 scripts/watchlist.py run all

```

When using the `--store` flag with the main command, results persist to `~/.config/last30days/briefings.db` (SQLite), enabling historical queries like `last30 what have you found about...`.

## Core File Architecture

| File | Responsibility |
|------|----------------|
| [`scripts/last30days.py`](https://github.com/mvanhorn/last30days-skill/blob/main/scripts/last30days.py) | CLI entry point and pipeline orchestration |
| [`scripts/lib/query_type.py`](https://github.com/mvanhorn/last30days-skill/blob/main/scripts/lib/query_type.py) | Regex classification and source tier rules |
| [`scripts/lib/score.py`](https://github.com/mvanhorn/last30days-skill/blob/main/scripts/lib/score.py) | Normalization and weighted scoring logic |
| [`scripts/lib/env.py`](https://github.com/mvanhorn/last30days-skill/blob/main/scripts/lib/env.py) | API key resolution and backend selection |
| [`scripts/lib/entity_extract.py`](https://github.com/mvanhorn/last30days-skill/blob/main/scripts/lib/entity_extract.py) | Phase 2 X handle and subreddit extraction |
| [`scripts/lib/dedupe.py`](https://github.com/mvanhorn/last30days-skill/blob/main/scripts/lib/dedupe.py) | URL-based duplicate removal |
| [`scripts/lib/render.py`](https://github.com/mvanhorn/last30days-skill/blob/main/scripts/lib/render.py) | Output formatting for all emit modes |
| [`scripts/lib/schema.py`](https://github.com/mvanhorn/last30days-skill/blob/main/scripts/lib/schema.py) | Pydantic models for type safety across sources |
| [`scripts/lib/reddit.py`](https://github.com/mvanhorn/last30days-skill/blob/main/scripts/lib/reddit.py) | ScrapeCreators/OpenAI-based Reddit search |
| [`scripts/lib/bird_x.py`](https://github.com/mvanhorn/last30days-skill/blob/main/scripts/lib/bird_x.py) | X/Twitter search via Bird GraphQL or xAI |
| [`scripts/watchlist.py`](https://github.com/mvanhorn/last30days-skill/blob/main/scripts/watchlist.py) | Scheduled research management |
| [`scripts/store.py`](https://github.com/mvanhorn/last30days-skill/blob/main/scripts/store.py) | SQLite persistence layer |

## Summary

- The last30days-skill runs as `python3 scripts/last30days.py "<query>"` and aggregates 30 days of social/web discussions
- Configuration loads from `.claude/last30days.env` or `~/.config/last30days/.env` via [`scripts/lib/env.py`](https://github.com/mvanhorn/last30days-skill/blob/main/scripts/lib/env.py)
- Source selection is query-type aware, using tiers defined in [`scripts/lib/query_type.py`](https://github.com/mvanhorn/last30days-skill/blob/main/scripts/lib/query_type.py) to minimize unnecessary API calls
- Scoring uses source-specific weights (e.g., 45/25/30 for Reddit/X) with penalties for web-search items to prioritize social signals
- Parallel execution with `ThreadPoolExecutor` and configurable `TIMEOUT_PROFILES` ensures responsive CLI performance
- Optional watchlist functionality in [`scripts/watchlist.py`](https://github.com/mvanhorn/last30days-skill/blob/main/scripts/watchlist.py) enables automated periodic research with SQLite storage

## Frequently Asked Questions

### What API keys are required to run the last30days skill?

You only need keys for the sources you intend to query. For basic social search, provide `SCRAPECREATORS_API_KEY` (Reddit) and `XAI_API_KEY` or Bird credentials (X). For web search, provide one of `PARALLEL_API_KEY`, `BRAVE_API_KEY`, or `OPENROUTER_API_KEY`. The [`env.py`](https://github.com/mvanhorn/last30days-skill/blob/main/env.py) module automatically detects which backends are available and skips sources with missing credentials unless explicitly requested.

### How does the supplemental search phase work?

After the initial search results are collected, [`entity_extract.py`](https://github.com/mvanhorn/last30days-skill/blob/main/entity_extract.py) scans content for X handles (e.g., `@username`) and subreddit names. The `_run_supplemental` function then executes focused searches for these entities in parallel. This surfaces high-engagement posts that discuss the topic without containing your original search keywords, improving recall for viral discussions and expert threads.

### Why do web search results have lower scores by default?

Web search results lack native engagement metrics (likes, comments), so the scoring algorithm in [`scripts/lib/score.py`](https://github.com/mvanhorn/last30days-skill/blob/main/scripts/lib/score.py) applies a `WEBSEARCH_SOURCE_PENALTY` (typically 15 points) to compensate. However, the penalty varies by query type—"concept" queries receive 0 penalty because authoritative documentation is preferred, while "product" or "prediction" queries apply the full penalty to prioritize social verification and sentiment.

### Can I add custom data sources to the pipeline?

Yes. Create a new module under `scripts/lib/` implementing `search_<source>` and `parse_<source>_response` functions that return typed items per [`scripts/lib/schema.py`](https://github.com/mvanhorn/last30days-skill/blob/main/scripts/lib/schema.py). Then register the source in the executor block within `run_research` in [`scripts/last30days.py`](https://github.com/mvanhorn/last30days-skill/blob/main/scripts/last30days.py) and add tier rules to `SOURCE_TIERS` in [`scripts/lib/query_type.py`](https://github.com/mvanhorn/last30days-skill/blob/main/scripts/lib/query_type.py) to control when it runs.