How to Load and Analyze Trade Data with Poly Data: A Complete Guide

Poly Data is a self-contained Python pipeline that fetches Polymarket order-filled events from the Goldsky subgraph, processes them into normalized CSV files, and exposes a clean Polars interface for quantitative analysis.

The warproxxx/poly_data repository provides a lightweight, idempotent data pipeline designed specifically for Polymarket trading data. Whether you're building trading strategies, conducting market research, or analyzing on-chain activity, this tool transforms raw blockchain events into tidy DataFrames ready for statistical analysis.

Understanding the Poly Data Pipeline Architecture

Poly Data follows a deliberate three-layer architecture that separates data collection, transformation, and analysis concerns. Each stage is idempotent—the pipeline counts existing rows and timestamps to resume automatically without duplicating data.

The workflow moves from raw API ingestion through Goldsky subgraph queries, ultimately landing in flat CSV files that support any analytics stack.

Step 1: Collecting Raw Trade Data from Polymarket

The collection layer consists of three independent modules that populate the underlying datasets required for analysis.

Fetching Market Metadata

update_utils/update_markets.py queries the Polymarket API to retrieve every available market definition, writing the results to markets.csv. This file serves as the canonical lookup table for market titles, resolutions, and token mappings.

Extracting Order-Filled Events

update_utils/update_goldsky.py connects to the Goldsky subgraph to pull raw orderFilled events, appending new rows to goldsky/orderFilled.csv. This module handles pagination and timestamp tracking to ensure incremental updates capture only new trading activity since the last run.

Discovering Missing Market Tokens

When trades reference tokens not present in the main markets file, poly_utils/utils.py provides the update_missing_tokens function. This utility fetches missing token IDs and stores them in missing_markets.csv, ensuring the pipeline maintains referential integrity even for newly created or obscure markets.

Step 2: Processing Raw Events into Analysis-Ready Trades

update_utils/process_live.py transforms the raw Goldsky events into a clean, denormalized trades table through several normalization steps:

  • Market joining: The get_markets() function loads both markets.csv and missing_markets.csv, de-duplicating by market id to return a unified Polars DataFrame
  • Asset identification: The script extracts the nonusdc_asset_id from each order (the asset ID that isn't "0")
  • Side derivation: Based on USDC positioning, it creates makerAsset, takerAsset, maker_direction, and taker_direction fields
  • Price calculation: Derives the effective price as USDC per outcome token
  • Amount normalization: Divides raw token amounts by 10⁶ to standardize decimal units

The get_processed_df function (lines 15-100) contains the core transformation logic, while process_live (lines 102-184) handles incremental processing and writes results to processed/trades.csv with proper header preservation.

Step 3: Orchestrating the Full Pipeline

Rather than running modules individually, update_all.py provides a single entry point that executes the complete workflow:

from update_utils.update_markets import update_markets
from update_utils.update_goldsky import update_goldsky
from update_utils.process_live import process_live

if __name__ == "__main__":
    update_markets()
    update_goldsky()
    process_live()

Running uv run python update_all.py performs a full end-to-end refresh, downloading new market definitions, fetching latest trades, and regenerating the processed CSV.

Loading and Analyzing Trade Data with Polars

Once the pipeline populates processed/trades.csv, you can load data efficiently using Polars' lazy evaluation. The poly_utils/utils.py module provides get_markets() to load market metadata, while pl.scan_csv() enables out-of-core processing for large datasets:

import polars as pl
from poly_utils import get_markets

# Load market definitions (merged markets + missing markets)

markets_df = get_markets()

# Load trades with lazy evaluation for memory efficiency

trades = (
    pl.scan_csv("processed/trades.csv")
    .with_columns(pl.col("timestamp").str.to_datetime())
    .collect(streaming=True)
)

# Filter specific user activity

USER_ADDRESS = "0x9d84ce0306f8551e02efef1680475fc0f1dc1344"
user_trades = trades.filter(pl.col("maker") == USER_ADDRESS)

print(user_trades.head())

Computing Aggregated Metrics

Calculate total USD volume per market using Polars' expression syntax:

usd_by_market = (
    trades.groupby("market_id")
    .agg(pl.sum("usd_amount").alias("total_usd"))
)

Converting to Pandas for Visualization

For compatibility with matplotlib or seaborn, materialize as a pandas DataFrame:

import pandas as pd
import matplotlib.pyplot as plt

# Convert to pandas for plotting

user_trades_pd = user_trades.to_pandas()

# Resample to daily volume

daily_volume = (
    user_trades_pd
    .set_index("timestamp")
    .resample("D")["usd_amount"]
    .sum()
    .fillna(0)
)

daily_volume.plot(kind="bar", title="Daily USD Trading Volume")
plt.show()

Extending Your Analysis

Because Poly Data stores everything in flat CSV files, you can integrate any analytics tool that supports CSV ingestion. Common extensions include:

  • DuckDB: Run SQL queries directly against processed/trades.csv without loading into memory
  • Time-series analysis: Use pl.datetime_range or pandas resampling to create daily, weekly, or hourly aggregates
  • Backtrader integration: The backtrader_plotting/ directory contains Bokeh utilities for interactive strategy visualization

The CSV-based architecture ensures your analysis code remains completely decoupled from data ingestion, preserving reproducibility across different environments.

Summary

  • Three-stage collection: update_markets.py, update_goldsky.py, and update_missing_tokens populate raw data files idempotently
  • Transformation layer: process_live.py normalizes raw events into processed/trades.csv with calculated prices and standardized amounts
  • Orchestration: Run update_all.py to execute the complete pipeline with a single command
  • Analysis interface: Use get_markets() from poly_utils/utils.py and pl.scan_csv() for memory-efficient Polars workflows
  • Storage format: Flat CSV files enable interoperability with pandas, SQL, DuckDB, or visualization libraries

Frequently Asked Questions

What data sources does Poly Data use?

Poly Data combines the Polymarket REST API for market metadata with the Goldsky subgraph for order-filled events. According to the warproxxx/poly_data source code, update_markets.py fetches from the Polymarket API while update_goldsky.py queries the Goldsky-hosted Polymarket subgraph for on-chain trading events.

How does the pipeline handle incremental updates?

Each collection module is designed to be idempotent. The scripts count existing rows and track timestamps, allowing them to resume from the last known state without duplicating data. You can safely run update_all.py multiple times; it will only fetch and process new data since the previous execution.

Can I use pandas instead of Polars for analysis?

Yes. While the repository examples favor Polars for performance—specifically using pl.scan_csv() for lazy evaluation—you can convert to pandas at any point using .to_pandas(). The CSV storage format means any tool that reads CSVs (pandas, R, Excel, DuckDB) works with the output files.

What is the schema of the final trades.csv file?

The processed/trades.csv file contains normalized trading data with derived fields including timestamp, market_id, maker, taker, makerAsset, takerAsset, maker_direction, taker_direction, price (USDC per token), and normalized usd_amount values. The process_live.py transformation ensures all amounts are divided by 10⁶ to represent standard decimal units rather than raw blockchain integers.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →