# How Polymarket Integration Works in last30days-skill: A Technical Deep Dive

> Discover how Polymarket integration elevates last30days-skill. Learn about parallel Gamma API fetches, hybrid scoring, and enhanced query processing for superior results.

- Repository: [Matt Van Horn/last30days-skill](https://github.com/mvanhorn/last30days-skill)
- Tags: deep-dive
- Published: 2026-03-25

---

**The last30days-skill integrates Polymarket by expanding user queries into multiple search strings, fetching data in parallel from the Gamma API, and scoring results using a hybrid text-similarity and market-quality formula before rendering them alongside other research sources.**

The **last30days-skill** (maintained in `mvanhorn/last30days-skill`) treats Polymarket as a first-class research source for prediction market data. When users ask questions containing the keyword `polymarket` or trigger the generic "prediction" query type, the system executes a multi-pass search algorithm implemented primarily in [`scripts/lib/polymarket.py`](https://github.com/mvanhorn/last30days-skill/blob/main/scripts/lib/polymarket.py).

## Query Expansion and Parallel Search

The integration begins with intelligent query generation and parallel API execution to maximize result coverage.

### Multi-String Query Generation

The `_expand_queries()` function in [`scripts/lib/polymarket.py`](https://github.com/mvanhorn/last30days-skill/blob/main/scripts/lib/polymarket.py) transforms a single user topic into up to six distinct search strings (lines 43-93). First, it strips common conversational prefixes like "last 7 days" or "what are people saying about" to isolate the **core subject**. Then it generates variations including the full topic, the core subject alone, and individual significant words (excluding single-character tokens and low-signal stop words). This expansion ensures the Gamma API search captures markets that might not match the exact user phrasing but contain relevant keywords.

### Parallel API Execution with ThreadPoolExecutor

Each expanded query runs against the Polymarket Gamma API (`https://gamma-api.polymarket.com/public-search`) using a `ThreadPoolExecutor` capped at **8 concurrent workers** (lines 54-66). The `_run_queries_parallel()` function submits `(query, page)` pairs to `_search_single_query()`, which performs HTTP GET requests via the shared `http` wrapper (lines 33-42). 

The search depth is controlled by `DEPTH_CONFIG` defined at the module level: `quick` mode fetches 1 page, `default` fetches 3 pages, and `deep` fetches 4 pages per query (lines 21-26). Results merge into a deduplicated dictionary keyed by event ID, where the earliest-matched query retains priority (lines 72-80). The final output respects `RESULT_CAP` to limit total events returned (lines 28-33).

## Domain-Level Expansion and Deduplication

After the initial pass, the engine performs a **second expansion phase** using domain-specific signals extracted from event tags. The `_extract_domain_queries()` function (lines 99-129) analyzes tags from all first-pass events, filtering against a generic-tag whitelist and excluding terms already present in the original query. The two most frequent non-generic tags appearing in at least two events become new search queries. These domain queries execute through `_run_queries_parallel()` once more (single page only) and merge into the existing event pool, capturing niche markets that semantic expansion might miss.

## Response Parsing and Relevance Scoring

The `parse_polymarket_response()` function (lines 80-130) transforms raw API payloads into normalized **PolymarketItem** objects suitable for the skill's generic rendering pipeline.

### Market Filtering and Selection

The parser first filters out closed or resolved events, retaining only active markets with verified liquidity (lines 100-125). For each event, it selects the **most liquid market** (sorted by `market_volume` descending) as the representative market (lines 29-36). It then collects outcome names from all active markets within the event, synthesizing binary outcomes when sub-markets provide team-specific probabilities beyond simple "Yes/No" options (lines 40-75).

### Hybrid Scoring Algorithm

Each item receives a relevance score calculated as:

```python
relevance = min(1.0, text_score * (0.75 + 0.25 * market_quality))

```

This formula (lines 34-35) ensures **text similarity dominates** (75% weight) while market quality signals provide a 25% boost or penalty. The `text_score` measures alignment between the core subject and market titles/outcomes. The `market_quality` component factors in volume, liquidity, price movement, and competitive bonuses. The parser also reorders outcome lists to display the topic-matching outcome first (lines 36-53) before applying the `_cap` limit supplied by the search function (lines 77-80).

## Integration with the Research Pipeline

The public entry point `search_polymarket()` is invoked from [`scripts/last30days.py`](https://github.com/mvanhorn/last30days-skill/blob/main/scripts/last30days.py) when the dispatcher detects a **prediction** query type via `detect_query_type()`. The dispatcher adds returned items to the research report under the `polymarket` key, where [`scripts/lib/render.py`](https://github.com/mvanhorn/last30days-skill/blob/main/scripts/lib/render.py) processes them identically to Reddit, X, or YouTube sources. The [`scripts/lib/normalize.py`](https://github.com/mvanhorn/last30days-skill/blob/main/scripts/lib/normalize.py) module further enriches these items with engagement metrics, while [`scripts/lib/score.py`](https://github.com/mvanhorn/last30days-skill/blob/main/scripts/lib/score.py) applies final relevance weighting across all sources.

## Implementation Example

You can invoke the Polymarket integration directly for testing or custom workflows:

```python
from lib import polymarket

# Search for prediction markets about Arizona basketball

result = polymarket.search_polymarket(
    topic="Arizona basketball",
    from_date="2026-02-24",
    to_date="2026-03-25",
    depth="default"  # Options: "quick", "default", "deep"

)

# Parse and display top results

items = polymarket.parse_polymarket_response(result, topic="Arizona basketball")
for item in items[:3]:
    print(f"{item['title']} → {item['url']}")
    print(f"Relevance: {item['relevance']:.2f}")

```

For command-line usage, the skill accepts natural language queries:

```bash
python3 scripts/last30days.py "OpenAI prediction market" --emit=compact

```

The `--emit=compact` flag returns a concise summary including any Polymarket results found during the research cycle.

## Summary

- **Query expansion** generates up to six search variations from a single user topic by extracting core subjects and significant words in [`scripts/lib/polymarket.py`](https://github.com/mvanhorn/last30days-skill/blob/main/scripts/lib/polymarket.py) (lines 43-93).
- **Parallel execution** uses an 8-worker `ThreadPoolExecutor` to query the Gamma API (`https://gamma-api.polymarket.com/public-search`) with configurable depth (1-4 pages).
- **Domain expansion** performs a second pass using high-frequency tags from initial results to discover niche markets (lines 99-129).
- **Relevance scoring** combines text similarity (75% weight) with market quality signals (25% weight) to rank active, liquid markets.
- **Pipeline integration** occurs through `search_polymarket()` called by [`scripts/last30days.py`](https://github.com/mvanhorn/last30days-skill/blob/main/scripts/last30days.py) when query type detection identifies prediction market requests.

## Frequently Asked Questions

### What API endpoint does last30days-skill use for Polymarket data?

The skill queries the public Gamma API at `https://gamma-api.polymarket.com/public-search`. This endpoint is accessed through the `_search_single_query()` helper function in [`scripts/lib/polymarket.py`](https://github.com/mvanhorn/last30days-skill/blob/main/scripts/lib/polymarket.py) (lines 33-42), which constructs parameterized GET requests and returns raw event data for parsing and scoring.

### How does the skill handle rate limiting or API failures during parallel searches?

The `_run_queries_parallel()` function collects errors into an `errors` list while continuing execution, ensuring partial results from successful queries remain available. The shared `http` wrapper (used in lines 33-42) likely implements retry logic and connection pooling, though the specific error handling depends on the wrapper's implementation in the `lib` package.

### Why does the relevance formula weight text similarity higher than market quality?

The scoring algorithm uses `relevance = min(1.0, text_score * (0.75 + 0.25 * market_quality))` (lines 34-35) to prioritize semantic relevance over raw liquidity. This prevents high-volume but off-topic markets from dominating results, ensuring users see prediction markets that actually address their query subject, with market quality serving as a tie-breaker or modest booster for equally relevant matches.

### Can I adjust how many pages the skill fetches from Polymarket?

Yes. The `depth` parameter in `search_polymarket()` accepts three modes defined in `DEPTH_CONFIG` (lines 21-26): `quick` (1 page), `default` (3 pages), and `deep` (4 pages). Additionally, `RESULT_CAP` (lines 28-33) limits the total number of events returned regardless of depth setting, preventing response bloat while maintaining search comprehensiveness.