How to Run the Poly Data Pipeline: A Complete ETL Guide for Polymarket Data

To run the Poly Data pipeline, execute python update_all.py after installing Python 3.8+ and dependencies to trigger the complete three-stage ETL workflow that extracts Polymarket markets, fetches on-chain events from Goldsky, and produces an analysis-ready trades CSV.

The Poly Data pipeline is a resumable ETL workflow maintained in the warproxxx/poly_data repository for ingesting prediction market data. It aggregates public market metadata from Polymarket's Gamma API with on-chain order book events from Goldsky's GraphQL endpoint, generating cleaned datasets suitable for quantitative analysis. Running the pipeline requires no API keys and automatically resumes from checkpoints on subsequent executions.

Understanding the Poly Data Pipeline Architecture

The pipeline follows a strict Extract-Transform-Load pattern orchestrated by update_all.py. Each stage handles a specific data domain and implements independent resumability logic to prevent duplicate processing.

Stage 1: Market Metadata Extraction

Located in update_utils/update_markets.py, this stage queries the Polymarket Gamma API at https://gamma-api.polymarket.com/markets. The update_markets() function implements pagination with resumable offsets by counting existing lines in markets.csv via the helper count_csv_lines. It handles HTTP 500/429 errors using exponential back-off and writes market attributes—including createdAt, id, question, and token IDs—to markets.csv and optionally missing_markets.csv.

Stage 2: On-Chain Event Extraction

The update_utils/update_goldsky.py module fetches orderFilledEvents from the Goldsky subgraph using a sticky-timestamp cursor pattern. The implementation loads the last timestamp from goldsky/cursor_state.json (or falls back to the final line of goldsky/orderFilled.csv), then constructs GraphQL queries with timestamp_gt or id_gt parameters to guarantee gap-free ingestion. Events are flattened from nested JSON using flatten-json, converted to a pandas DataFrame, and appended to goldsky/orderFilled.csv before persisting the new cursor.

Stage 3: Trade Processing and Enrichment

The update_utils/process_live.py utility transforms raw events into analyst-ready data. It reads both markets.csv and missing_markets.csv via poly_utils/utils.get_markets into a Polars DataFrame keyed by market ID, then lazily loads the Goldsky events. The get_processed_df function identifies the non-USDC asset in each trade, calculates trade direction and USD amounts, and deduplicates records against existing processed/trades.csv entries before appending new rows.

Prerequisites and Installation

The pipeline requires Python 3.8 or higher and several scientific Python libraries. No API credentials are necessary, as all endpoints are public.

Clone the repository and install dependencies from pyproject.toml:

git clone https://github.com/warproxxx/poly_data.git
cd poly_data
python -m venv .venv
source .venv/bin/activate
pip install -e .

Core dependencies include pandas, polars, requests, gql[requests], and flatten-json. Optional Jupyter packages support the provided analysis notebooks.

Running the Full Pipeline

Execute the orchestrator script from the repository root:

python update_all.py

This sequential invocation calls update_markets(), update_goldsky(), and process_live(), producing three key artifacts:

  • markets.csv — Complete Polymarket market definitions with outcome tokens
  • goldsky/orderFilled.csv — Raw on-chain order fill events from the subgraph
  • processed/trades.csv — Enriched trade data joined with market metadata

The console displays progress messages for each stage. Because each component resumes automatically from the last processed timestamp or offset, re-running the command safely appends only new data without duplication.

Running Individual Pipeline Stages

For incremental updates or debugging specific components, execute individual stages directly via Python's -c flag:

Fetch only new market definitions:

python -c "from update_utils.update_markets import update_markets; update_markets()"

Pull only new on-chain events:

python -c "from update_utils.update_goldsky import update_goldsky; update_goldsky()"

Process pending trades without re-fetching source data:

python -c "from update_utils.process_live import process_live; process_live()"

Working with the Output Data

After the pipeline completes, load the processed trades into pandas for immediate analysis:

import pandas as pd

# Load the enriched dataset

trades = pd.read_csv("processed/trades.csv")
print(f"Loaded {len(trades)} trades")

# Calculate total USD volume per market

volume = trades.groupby("market_id")["usd_amount"].sum().sort_values(ascending=False)
print(volume.head())

The trades.csv schema includes calculated fields for trade direction, normalized token amounts, and human-readable asset identifiers derived from the market metadata join performed in process_live.py.

Summary

  • Poly Data is a resumable three-stage ETL pipeline for Polymarket data orchestrated by update_all.py in the warproxxx/poly_data repository.
  • Market extraction (update_markets.py) uses offset-based pagination with line-counting resume logic against the Gamma API.
  • Goldsky extraction (update_goldsky.py) implements cursor-based resumability using cursor_state.json or the last CSV line to prevent gaps in on-chain event history.
  • Processing (process_live.py) joins raw events with market metadata using Polars DataFrames to generate processed/trades.csv.
  • Python 3.8+ and standard data science dependencies are the only requirements; no API keys are necessary.

Frequently Asked Questions

What Python version is required to run the Poly Data pipeline?

Python 3.8 or higher is required. The codebase utilizes modern pandas and polars APIs that depend on contemporary Python features, though it remains compatible with standard CPython distributions and virtual environments.

Does the Poly Data pipeline require API keys for Polymarket or Goldsky?

No API keys are required. The pipeline accesses public endpoints only: the Polymarket Gamma API for market metadata and the Goldsky public GraphQL endpoint for on-chain events. Built-in exponential back-off handles rate limiting automatically.

How does the pipeline handle interruptions or crashes during execution?

Each stage implements independent resumability. Market extraction tracks the current offset by counting lines in markets.csv via count_csv_lines, while Goldsky extraction persists a timestamp cursor to goldsky/cursor_state.json after each batch. Processed trade deduplication ensures appending only new records to processed/trades.csv, making the pipeline safe to restart at any point.

What distinguishes the raw Goldsky output from the processed trades CSV?

The goldsky/orderFilled.csv contains raw, flattened GraphQL event data with blockchain-specific identifiers, maker/taker addresses, and epoch timestamps. The processed/trades.csv enriches these records with human-readable market questions from markets.csv, calculated USD values, trade directions (buy/sell), and normalized timestamps suitable for quantitative analysis and backtesting.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →