# How to Run the Poly Data Pipeline: A Complete ETL Guide for Polymarket Data

> Run the Poly Data pipeline effortlessly. Execute python update_all.py to complete the Polymarket ETL workflow, extracting market data and on-chain events for analysis.

- Repository: [warproxxx/poly_data](https://github.com/warproxxx/poly_data)
- Tags: how-to-guide
- Published: 2026-04-21

---

**To run the Poly Data pipeline, execute `python update_all.py` after installing Python 3.8+ and dependencies to trigger the complete three-stage ETL workflow that extracts Polymarket markets, fetches on-chain events from Goldsky, and produces an analysis-ready trades CSV.**

The **Poly Data** pipeline is a resumable ETL workflow maintained in the `warproxxx/poly_data` repository for ingesting prediction market data. It aggregates public market metadata from Polymarket's Gamma API with on-chain order book events from Goldsky's GraphQL endpoint, generating cleaned datasets suitable for quantitative analysis. Running the pipeline requires no API keys and automatically resumes from checkpoints on subsequent executions.

## Understanding the Poly Data Pipeline Architecture

The pipeline follows a strict Extract-Transform-Load pattern orchestrated by [`update_all.py`](https://github.com/warproxxx/poly_data/blob/main/update_all.py). Each stage handles a specific data domain and implements independent resumability logic to prevent duplicate processing.

### Stage 1: Market Metadata Extraction

Located in [`update_utils/update_markets.py`](https://github.com/warproxxx/poly_data/blob/main/update_utils/update_markets.py), this stage queries the Polymarket Gamma API at `https://gamma-api.polymarket.com/markets`. The `update_markets()` function implements pagination with resumable offsets by counting existing lines in `markets.csv` via the helper `count_csv_lines`. It handles HTTP 500/429 errors using exponential back-off and writes market attributes—including `createdAt`, `id`, `question`, and token IDs—to `markets.csv` and optionally `missing_markets.csv`.

### Stage 2: On-Chain Event Extraction

The [`update_utils/update_goldsky.py`](https://github.com/warproxxx/poly_data/blob/main/update_utils/update_goldsky.py) module fetches `orderFilledEvents` from the Goldsky subgraph using a sticky-timestamp cursor pattern. The implementation loads the last timestamp from [`goldsky/cursor_state.json`](https://github.com/warproxxx/poly_data/blob/main/goldsky/cursor_state.json) (or falls back to the final line of `goldsky/orderFilled.csv`), then constructs GraphQL queries with `timestamp_gt` or `id_gt` parameters to guarantee gap-free ingestion. Events are flattened from nested JSON using `flatten-json`, converted to a **pandas** DataFrame, and appended to `goldsky/orderFilled.csv` before persisting the new cursor.

### Stage 3: Trade Processing and Enrichment

The [`update_utils/process_live.py`](https://github.com/warproxxx/poly_data/blob/main/update_utils/process_live.py) utility transforms raw events into analyst-ready data. It reads both `markets.csv` and `missing_markets.csv` via `poly_utils/utils.get_markets` into a **Polars** DataFrame keyed by market ID, then lazily loads the Goldsky events. The `get_processed_df` function identifies the non-USDC asset in each trade, calculates trade direction and USD amounts, and deduplicates records against existing `processed/trades.csv` entries before appending new rows.

## Prerequisites and Installation

The pipeline requires Python 3.8 or higher and several scientific Python libraries. No API credentials are necessary, as all endpoints are public.

Clone the repository and install dependencies from [`pyproject.toml`](https://github.com/warproxxx/poly_data/blob/main/pyproject.toml):

```bash
git clone https://github.com/warproxxx/poly_data.git
cd poly_data
python -m venv .venv
source .venv/bin/activate
pip install -e .

```

Core dependencies include **pandas**, **polars**, **requests**, **gql[requests]**, and **flatten-json**. Optional Jupyter packages support the provided analysis notebooks.

## Running the Full Pipeline

Execute the orchestrator script from the repository root:

```bash
python update_all.py

```

This sequential invocation calls `update_markets()`, `update_goldsky()`, and `process_live()`, producing three key artifacts:
- `markets.csv` — Complete Polymarket market definitions with outcome tokens
- `goldsky/orderFilled.csv` — Raw on-chain order fill events from the subgraph
- `processed/trades.csv` — Enriched trade data joined with market metadata

The console displays progress messages for each stage. Because each component resumes automatically from the last processed timestamp or offset, re-running the command safely appends only new data without duplication.

## Running Individual Pipeline Stages

For incremental updates or debugging specific components, execute individual stages directly via Python's `-c` flag:

Fetch only new market definitions:

```bash
python -c "from update_utils.update_markets import update_markets; update_markets()"

```

Pull only new on-chain events:

```bash
python -c "from update_utils.update_goldsky import update_goldsky; update_goldsky()"

```

Process pending trades without re-fetching source data:

```bash
python -c "from update_utils.process_live import process_live; process_live()"

```

## Working with the Output Data

After the pipeline completes, load the processed trades into **pandas** for immediate analysis:

```python
import pandas as pd

# Load the enriched dataset

trades = pd.read_csv("processed/trades.csv")
print(f"Loaded {len(trades)} trades")

# Calculate total USD volume per market

volume = trades.groupby("market_id")["usd_amount"].sum().sort_values(ascending=False)
print(volume.head())

```

The `trades.csv` schema includes calculated fields for trade direction, normalized token amounts, and human-readable asset identifiers derived from the market metadata join performed in [`process_live.py`](https://github.com/warproxxx/poly_data/blob/main/process_live.py).

## Summary

- **Poly Data** is a resumable three-stage ETL pipeline for Polymarket data orchestrated by [`update_all.py`](https://github.com/warproxxx/poly_data/blob/main/update_all.py) in the `warproxxx/poly_data` repository.
- **Market extraction** ([`update_markets.py`](https://github.com/warproxxx/poly_data/blob/main/update_markets.py)) uses offset-based pagination with line-counting resume logic against the Gamma API.
- **Goldsky extraction** ([`update_goldsky.py`](https://github.com/warproxxx/poly_data/blob/main/update_goldsky.py)) implements cursor-based resumability using [`cursor_state.json`](https://github.com/warproxxx/poly_data/blob/main/cursor_state.json) or the last CSV line to prevent gaps in on-chain event history.
- **Processing** ([`process_live.py`](https://github.com/warproxxx/poly_data/blob/main/process_live.py)) joins raw events with market metadata using Polars DataFrames to generate `processed/trades.csv`.
- Python 3.8+ and standard data science dependencies are the only requirements; no API keys are necessary.

## Frequently Asked Questions

### What Python version is required to run the Poly Data pipeline?

Python 3.8 or higher is required. The codebase utilizes modern **pandas** and **polars** APIs that depend on contemporary Python features, though it remains compatible with standard CPython distributions and virtual environments.

### Does the Poly Data pipeline require API keys for Polymarket or Goldsky?

No API keys are required. The pipeline accesses public endpoints only: the Polymarket Gamma API for market metadata and the Goldsky public GraphQL endpoint for on-chain events. Built-in exponential back-off handles rate limiting automatically.

### How does the pipeline handle interruptions or crashes during execution?

Each stage implements independent resumability. Market extraction tracks the current offset by counting lines in `markets.csv` via `count_csv_lines`, while Goldsky extraction persists a timestamp cursor to [`goldsky/cursor_state.json`](https://github.com/warproxxx/poly_data/blob/main/goldsky/cursor_state.json) after each batch. Processed trade deduplication ensures appending only new records to `processed/trades.csv`, making the pipeline safe to restart at any point.

### What distinguishes the raw Goldsky output from the processed trades CSV?

The `goldsky/orderFilled.csv` contains raw, flattened GraphQL event data with blockchain-specific identifiers, maker/taker addresses, and epoch timestamps. The `processed/trades.csv` enriches these records with human-readable market questions from `markets.csv`, calculated USD values, trade directions (buy/sell), and normalized timestamps suitable for quantitative analysis and backtesting.