How Polymarket Integration Works in last30days-skill: A Technical Deep Dive

The last30days-skill integrates Polymarket by expanding user queries into multiple search strings, fetching data in parallel from the Gamma API, and scoring results using a hybrid text-similarity and market-quality formula before rendering them alongside other research sources.

The last30days-skill (maintained in mvanhorn/last30days-skill) treats Polymarket as a first-class research source for prediction market data. When users ask questions containing the keyword polymarket or trigger the generic "prediction" query type, the system executes a multi-pass search algorithm implemented primarily in scripts/lib/polymarket.py.

The integration begins with intelligent query generation and parallel API execution to maximize result coverage.

Multi-String Query Generation

The _expand_queries() function in scripts/lib/polymarket.py transforms a single user topic into up to six distinct search strings (lines 43-93). First, it strips common conversational prefixes like "last 7 days" or "what are people saying about" to isolate the core subject. Then it generates variations including the full topic, the core subject alone, and individual significant words (excluding single-character tokens and low-signal stop words). This expansion ensures the Gamma API search captures markets that might not match the exact user phrasing but contain relevant keywords.

Parallel API Execution with ThreadPoolExecutor

Each expanded query runs against the Polymarket Gamma API (https://gamma-api.polymarket.com/public-search) using a ThreadPoolExecutor capped at 8 concurrent workers (lines 54-66). The _run_queries_parallel() function submits (query, page) pairs to _search_single_query(), which performs HTTP GET requests via the shared http wrapper (lines 33-42).

The search depth is controlled by DEPTH_CONFIG defined at the module level: quick mode fetches 1 page, default fetches 3 pages, and deep fetches 4 pages per query (lines 21-26). Results merge into a deduplicated dictionary keyed by event ID, where the earliest-matched query retains priority (lines 72-80). The final output respects RESULT_CAP to limit total events returned (lines 28-33).

Domain-Level Expansion and Deduplication

After the initial pass, the engine performs a second expansion phase using domain-specific signals extracted from event tags. The _extract_domain_queries() function (lines 99-129) analyzes tags from all first-pass events, filtering against a generic-tag whitelist and excluding terms already present in the original query. The two most frequent non-generic tags appearing in at least two events become new search queries. These domain queries execute through _run_queries_parallel() once more (single page only) and merge into the existing event pool, capturing niche markets that semantic expansion might miss.

Response Parsing and Relevance Scoring

The parse_polymarket_response() function (lines 80-130) transforms raw API payloads into normalized PolymarketItem objects suitable for the skill's generic rendering pipeline.

Market Filtering and Selection

The parser first filters out closed or resolved events, retaining only active markets with verified liquidity (lines 100-125). For each event, it selects the most liquid market (sorted by market_volume descending) as the representative market (lines 29-36). It then collects outcome names from all active markets within the event, synthesizing binary outcomes when sub-markets provide team-specific probabilities beyond simple "Yes/No" options (lines 40-75).

Hybrid Scoring Algorithm

Each item receives a relevance score calculated as:

relevance = min(1.0, text_score * (0.75 + 0.25 * market_quality))

This formula (lines 34-35) ensures text similarity dominates (75% weight) while market quality signals provide a 25% boost or penalty. The text_score measures alignment between the core subject and market titles/outcomes. The market_quality component factors in volume, liquidity, price movement, and competitive bonuses. The parser also reorders outcome lists to display the topic-matching outcome first (lines 36-53) before applying the _cap limit supplied by the search function (lines 77-80).

Integration with the Research Pipeline

The public entry point search_polymarket() is invoked from scripts/last30days.py when the dispatcher detects a prediction query type via detect_query_type(). The dispatcher adds returned items to the research report under the polymarket key, where scripts/lib/render.py processes them identically to Reddit, X, or YouTube sources. The scripts/lib/normalize.py module further enriches these items with engagement metrics, while scripts/lib/score.py applies final relevance weighting across all sources.

Implementation Example

You can invoke the Polymarket integration directly for testing or custom workflows:

from lib import polymarket

# Search for prediction markets about Arizona basketball

result = polymarket.search_polymarket(
    topic="Arizona basketball",
    from_date="2026-02-24",
    to_date="2026-03-25",
    depth="default"  # Options: "quick", "default", "deep"

)

# Parse and display top results

items = polymarket.parse_polymarket_response(result, topic="Arizona basketball")
for item in items[:3]:
    print(f"{item['title']} → {item['url']}")
    print(f"Relevance: {item['relevance']:.2f}")

For command-line usage, the skill accepts natural language queries:

python3 scripts/last30days.py "OpenAI prediction market" --emit=compact

The --emit=compact flag returns a concise summary including any Polymarket results found during the research cycle.

Summary

  • Query expansion generates up to six search variations from a single user topic by extracting core subjects and significant words in scripts/lib/polymarket.py (lines 43-93).
  • Parallel execution uses an 8-worker ThreadPoolExecutor to query the Gamma API (https://gamma-api.polymarket.com/public-search) with configurable depth (1-4 pages).
  • Domain expansion performs a second pass using high-frequency tags from initial results to discover niche markets (lines 99-129).
  • Relevance scoring combines text similarity (75% weight) with market quality signals (25% weight) to rank active, liquid markets.
  • Pipeline integration occurs through search_polymarket() called by scripts/last30days.py when query type detection identifies prediction market requests.

Frequently Asked Questions

What API endpoint does last30days-skill use for Polymarket data?

The skill queries the public Gamma API at https://gamma-api.polymarket.com/public-search. This endpoint is accessed through the _search_single_query() helper function in scripts/lib/polymarket.py (lines 33-42), which constructs parameterized GET requests and returns raw event data for parsing and scoring.

How does the skill handle rate limiting or API failures during parallel searches?

The _run_queries_parallel() function collects errors into an errors list while continuing execution, ensuring partial results from successful queries remain available. The shared http wrapper (used in lines 33-42) likely implements retry logic and connection pooling, though the specific error handling depends on the wrapper's implementation in the lib package.

Why does the relevance formula weight text similarity higher than market quality?

The scoring algorithm uses relevance = min(1.0, text_score * (0.75 + 0.25 * market_quality)) (lines 34-35) to prioritize semantic relevance over raw liquidity. This prevents high-volume but off-topic markets from dominating results, ensuring users see prediction markets that actually address their query subject, with market quality serving as a tie-breaker or modest booster for equally relevant matches.

Can I adjust how many pages the skill fetches from Polymarket?

Yes. The depth parameter in search_polymarket() accepts three modes defined in DEPTH_CONFIG (lines 21-26): quick (1 page), default (3 pages), and deep (4 pages). Additionally, RESULT_CAP (lines 28-33) limits the total number of events returned regardless of depth setting, preventing response bloat while maintaining search comprehensiveness.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →