# How to Load and Analyze Trade Data with Poly Data: A Complete Guide

> Learn to load and analyze trade data with Poly Data. This Python pipeline processes Polymarket events into normalized CSVs for quantitative analysis using Polars.

- Repository: [warproxxx/poly_data](https://github.com/warproxxx/poly_data)
- Tags: how-to-guide
- Published: 2026-04-21

---

**Poly Data is a self-contained Python pipeline that fetches Polymarket order-filled events from the Goldsky subgraph, processes them into normalized CSV files, and exposes a clean Polars interface for quantitative analysis.**

The `warproxxx/poly_data` repository provides a lightweight, idempotent data pipeline designed specifically for Polymarket trading data. Whether you're building trading strategies, conducting market research, or analyzing on-chain activity, this tool transforms raw blockchain events into tidy DataFrames ready for statistical analysis.

## Understanding the Poly Data Pipeline Architecture

Poly Data follows a deliberate three-layer architecture that separates data collection, transformation, and analysis concerns. Each stage is **idempotent**—the pipeline counts existing rows and timestamps to resume automatically without duplicating data.

The workflow moves from raw API ingestion through Goldsky subgraph queries, ultimately landing in flat CSV files that support any analytics stack.

## Step 1: Collecting Raw Trade Data from Polymarket

The collection layer consists of three independent modules that populate the underlying datasets required for analysis.

### Fetching Market Metadata

[`update_utils/update_markets.py`](https://github.com/warproxxx/poly_data/blob/main/update_utils/update_markets.py) queries the Polymarket API to retrieve every available market definition, writing the results to `markets.csv`. This file serves as the canonical lookup table for market titles, resolutions, and token mappings.

### Extracting Order-Filled Events

[`update_utils/update_goldsky.py`](https://github.com/warproxxx/poly_data/blob/main/update_utils/update_goldsky.py) connects to the Goldsky subgraph to pull raw `orderFilled` events, appending new rows to `goldsky/orderFilled.csv`. This module handles pagination and timestamp tracking to ensure incremental updates capture only new trading activity since the last run.

### Discovering Missing Market Tokens

When trades reference tokens not present in the main markets file, [`poly_utils/utils.py`](https://github.com/warproxxx/poly_data/blob/main/poly_utils/utils.py) provides the `update_missing_tokens` function. This utility fetches missing token IDs and stores them in `missing_markets.csv`, ensuring the pipeline maintains referential integrity even for newly created or obscure markets.

## Step 2: Processing Raw Events into Analysis-Ready Trades

[`update_utils/process_live.py`](https://github.com/warproxxx/poly_data/blob/main/update_utils/process_live.py) transforms the raw Goldsky events into a clean, denormalized trades table through several normalization steps:

- **Market joining**: The `get_markets()` function loads both `markets.csv` and `missing_markets.csv`, de-duplicating by market `id` to return a unified Polars DataFrame
- **Asset identification**: The script extracts the `nonusdc_asset_id` from each order (the asset ID that isn't `"0"`)
- **Side derivation**: Based on USDC positioning, it creates `makerAsset`, `takerAsset`, `maker_direction`, and `taker_direction` fields
- **Price calculation**: Derives the effective price as USDC per outcome token
- **Amount normalization**: Divides raw token amounts by `10⁶` to standardize decimal units

The `get_processed_df` function (lines 15-100) contains the core transformation logic, while `process_live` (lines 102-184) handles incremental processing and writes results to `processed/trades.csv` with proper header preservation.

## Step 3: Orchestrating the Full Pipeline

Rather than running modules individually, [`update_all.py`](https://github.com/warproxxx/poly_data/blob/main/update_all.py) provides a single entry point that executes the complete workflow:

```python
from update_utils.update_markets import update_markets
from update_utils.update_goldsky import update_goldsky
from update_utils.process_live import process_live

if __name__ == "__main__":
    update_markets()
    update_goldsky()
    process_live()

```

Running `uv run python update_all.py` performs a full end-to-end refresh, downloading new market definitions, fetching latest trades, and regenerating the processed CSV.

## Loading and Analyzing Trade Data with Polars

Once the pipeline populates `processed/trades.csv`, you can load data efficiently using Polars' lazy evaluation. The [`poly_utils/utils.py`](https://github.com/warproxxx/poly_data/blob/main/poly_utils/utils.py) module provides `get_markets()` to load market metadata, while `pl.scan_csv()` enables out-of-core processing for large datasets:

```python
import polars as pl
from poly_utils import get_markets

# Load market definitions (merged markets + missing markets)

markets_df = get_markets()

# Load trades with lazy evaluation for memory efficiency

trades = (
    pl.scan_csv("processed/trades.csv")
    .with_columns(pl.col("timestamp").str.to_datetime())
    .collect(streaming=True)
)

# Filter specific user activity

USER_ADDRESS = "0x9d84ce0306f8551e02efef1680475fc0f1dc1344"
user_trades = trades.filter(pl.col("maker") == USER_ADDRESS)

print(user_trades.head())

```

### Computing Aggregated Metrics

Calculate total USD volume per market using Polars' expression syntax:

```python
usd_by_market = (
    trades.groupby("market_id")
    .agg(pl.sum("usd_amount").alias("total_usd"))
)

```

### Converting to Pandas for Visualization

For compatibility with matplotlib or seaborn, materialize as a pandas DataFrame:

```python
import pandas as pd
import matplotlib.pyplot as plt

# Convert to pandas for plotting

user_trades_pd = user_trades.to_pandas()

# Resample to daily volume

daily_volume = (
    user_trades_pd
    .set_index("timestamp")
    .resample("D")["usd_amount"]
    .sum()
    .fillna(0)
)

daily_volume.plot(kind="bar", title="Daily USD Trading Volume")
plt.show()

```

## Extending Your Analysis

Because Poly Data stores everything in **flat CSV files**, you can integrate any analytics tool that supports CSV ingestion. Common extensions include:

- **DuckDB**: Run SQL queries directly against `processed/trades.csv` without loading into memory
- **Time-series analysis**: Use `pl.datetime_range` or pandas resampling to create daily, weekly, or hourly aggregates
- **Backtrader integration**: The `backtrader_plotting/` directory contains Bokeh utilities for interactive strategy visualization

The CSV-based architecture ensures your analysis code remains completely decoupled from data ingestion, preserving reproducibility across different environments.

## Summary

- **Three-stage collection**: [`update_markets.py`](https://github.com/warproxxx/poly_data/blob/main/update_markets.py), [`update_goldsky.py`](https://github.com/warproxxx/poly_data/blob/main/update_goldsky.py), and `update_missing_tokens` populate raw data files idempotently
- **Transformation layer**: [`process_live.py`](https://github.com/warproxxx/poly_data/blob/main/process_live.py) normalizes raw events into `processed/trades.csv` with calculated prices and standardized amounts
- **Orchestration**: Run [`update_all.py`](https://github.com/warproxxx/poly_data/blob/main/update_all.py) to execute the complete pipeline with a single command
- **Analysis interface**: Use `get_markets()` from [`poly_utils/utils.py`](https://github.com/warproxxx/poly_data/blob/main/poly_utils/utils.py) and `pl.scan_csv()` for memory-efficient Polars workflows
- **Storage format**: Flat CSV files enable interoperability with pandas, SQL, DuckDB, or visualization libraries

## Frequently Asked Questions

### What data sources does Poly Data use?

Poly Data combines the Polymarket REST API for market metadata with the Goldsky subgraph for order-filled events. According to the `warproxxx/poly_data` source code, [`update_markets.py`](https://github.com/warproxxx/poly_data/blob/main/update_markets.py) fetches from the Polymarket API while [`update_goldsky.py`](https://github.com/warproxxx/poly_data/blob/main/update_goldsky.py) queries the Goldsky-hosted Polymarket subgraph for on-chain trading events.

### How does the pipeline handle incremental updates?

Each collection module is designed to be idempotent. The scripts count existing rows and track timestamps, allowing them to resume from the last known state without duplicating data. You can safely run [`update_all.py`](https://github.com/warproxxx/poly_data/blob/main/update_all.py) multiple times; it will only fetch and process new data since the previous execution.

### Can I use pandas instead of Polars for analysis?

Yes. While the repository examples favor Polars for performance—specifically using `pl.scan_csv()` for lazy evaluation—you can convert to pandas at any point using `.to_pandas()`. The CSV storage format means any tool that reads CSVs (pandas, R, Excel, DuckDB) works with the output files.

### What is the schema of the final trades.csv file?

The `processed/trades.csv` file contains normalized trading data with derived fields including `timestamp`, `market_id`, `maker`, `taker`, `makerAsset`, `takerAsset`, `maker_direction`, `taker_direction`, `price` (USDC per token), and normalized `usd_amount` values. The [`process_live.py`](https://github.com/warproxxx/poly_data/blob/main/process_live.py) transformation ensures all amounts are divided by `10⁶` to represent standard decimal units rather than raw blockchain integers.