How to Load and Analyze Trade Data with Poly Data: A Complete Guide
Poly Data is a self-contained Python pipeline that fetches Polymarket order-filled events from the Goldsky subgraph, processes them into normalized CSV files, and exposes a clean Polars interface for quantitative analysis.
The warproxxx/poly_data repository provides a lightweight, idempotent data pipeline designed specifically for Polymarket trading data. Whether you're building trading strategies, conducting market research, or analyzing on-chain activity, this tool transforms raw blockchain events into tidy DataFrames ready for statistical analysis.
Understanding the Poly Data Pipeline Architecture
Poly Data follows a deliberate three-layer architecture that separates data collection, transformation, and analysis concerns. Each stage is idempotent—the pipeline counts existing rows and timestamps to resume automatically without duplicating data.
The workflow moves from raw API ingestion through Goldsky subgraph queries, ultimately landing in flat CSV files that support any analytics stack.
Step 1: Collecting Raw Trade Data from Polymarket
The collection layer consists of three independent modules that populate the underlying datasets required for analysis.
Fetching Market Metadata
update_utils/update_markets.py queries the Polymarket API to retrieve every available market definition, writing the results to markets.csv. This file serves as the canonical lookup table for market titles, resolutions, and token mappings.
Extracting Order-Filled Events
update_utils/update_goldsky.py connects to the Goldsky subgraph to pull raw orderFilled events, appending new rows to goldsky/orderFilled.csv. This module handles pagination and timestamp tracking to ensure incremental updates capture only new trading activity since the last run.
Discovering Missing Market Tokens
When trades reference tokens not present in the main markets file, poly_utils/utils.py provides the update_missing_tokens function. This utility fetches missing token IDs and stores them in missing_markets.csv, ensuring the pipeline maintains referential integrity even for newly created or obscure markets.
Step 2: Processing Raw Events into Analysis-Ready Trades
update_utils/process_live.py transforms the raw Goldsky events into a clean, denormalized trades table through several normalization steps:
- Market joining: The
get_markets()function loads bothmarkets.csvandmissing_markets.csv, de-duplicating by marketidto return a unified Polars DataFrame - Asset identification: The script extracts the
nonusdc_asset_idfrom each order (the asset ID that isn't"0") - Side derivation: Based on USDC positioning, it creates
makerAsset,takerAsset,maker_direction, andtaker_directionfields - Price calculation: Derives the effective price as USDC per outcome token
- Amount normalization: Divides raw token amounts by
10⁶to standardize decimal units
The get_processed_df function (lines 15-100) contains the core transformation logic, while process_live (lines 102-184) handles incremental processing and writes results to processed/trades.csv with proper header preservation.
Step 3: Orchestrating the Full Pipeline
Rather than running modules individually, update_all.py provides a single entry point that executes the complete workflow:
from update_utils.update_markets import update_markets
from update_utils.update_goldsky import update_goldsky
from update_utils.process_live import process_live
if __name__ == "__main__":
update_markets()
update_goldsky()
process_live()
Running uv run python update_all.py performs a full end-to-end refresh, downloading new market definitions, fetching latest trades, and regenerating the processed CSV.
Loading and Analyzing Trade Data with Polars
Once the pipeline populates processed/trades.csv, you can load data efficiently using Polars' lazy evaluation. The poly_utils/utils.py module provides get_markets() to load market metadata, while pl.scan_csv() enables out-of-core processing for large datasets:
import polars as pl
from poly_utils import get_markets
# Load market definitions (merged markets + missing markets)
markets_df = get_markets()
# Load trades with lazy evaluation for memory efficiency
trades = (
pl.scan_csv("processed/trades.csv")
.with_columns(pl.col("timestamp").str.to_datetime())
.collect(streaming=True)
)
# Filter specific user activity
USER_ADDRESS = "0x9d84ce0306f8551e02efef1680475fc0f1dc1344"
user_trades = trades.filter(pl.col("maker") == USER_ADDRESS)
print(user_trades.head())
Computing Aggregated Metrics
Calculate total USD volume per market using Polars' expression syntax:
usd_by_market = (
trades.groupby("market_id")
.agg(pl.sum("usd_amount").alias("total_usd"))
)
Converting to Pandas for Visualization
For compatibility with matplotlib or seaborn, materialize as a pandas DataFrame:
import pandas as pd
import matplotlib.pyplot as plt
# Convert to pandas for plotting
user_trades_pd = user_trades.to_pandas()
# Resample to daily volume
daily_volume = (
user_trades_pd
.set_index("timestamp")
.resample("D")["usd_amount"]
.sum()
.fillna(0)
)
daily_volume.plot(kind="bar", title="Daily USD Trading Volume")
plt.show()
Extending Your Analysis
Because Poly Data stores everything in flat CSV files, you can integrate any analytics tool that supports CSV ingestion. Common extensions include:
- DuckDB: Run SQL queries directly against
processed/trades.csvwithout loading into memory - Time-series analysis: Use
pl.datetime_rangeor pandas resampling to create daily, weekly, or hourly aggregates - Backtrader integration: The
backtrader_plotting/directory contains Bokeh utilities for interactive strategy visualization
The CSV-based architecture ensures your analysis code remains completely decoupled from data ingestion, preserving reproducibility across different environments.
Summary
- Three-stage collection:
update_markets.py,update_goldsky.py, andupdate_missing_tokenspopulate raw data files idempotently - Transformation layer:
process_live.pynormalizes raw events intoprocessed/trades.csvwith calculated prices and standardized amounts - Orchestration: Run
update_all.pyto execute the complete pipeline with a single command - Analysis interface: Use
get_markets()frompoly_utils/utils.pyandpl.scan_csv()for memory-efficient Polars workflows - Storage format: Flat CSV files enable interoperability with pandas, SQL, DuckDB, or visualization libraries
Frequently Asked Questions
What data sources does Poly Data use?
Poly Data combines the Polymarket REST API for market metadata with the Goldsky subgraph for order-filled events. According to the warproxxx/poly_data source code, update_markets.py fetches from the Polymarket API while update_goldsky.py queries the Goldsky-hosted Polymarket subgraph for on-chain trading events.
How does the pipeline handle incremental updates?
Each collection module is designed to be idempotent. The scripts count existing rows and track timestamps, allowing them to resume from the last known state without duplicating data. You can safely run update_all.py multiple times; it will only fetch and process new data since the previous execution.
Can I use pandas instead of Polars for analysis?
Yes. While the repository examples favor Polars for performance—specifically using pl.scan_csv() for lazy evaluation—you can convert to pandas at any point using .to_pandas(). The CSV storage format means any tool that reads CSVs (pandas, R, Excel, DuckDB) works with the output files.
What is the schema of the final trades.csv file?
The processed/trades.csv file contains normalized trading data with derived fields including timestamp, market_id, maker, taker, makerAsset, takerAsset, maker_direction, taker_direction, price (USDC per token), and normalized usd_amount values. The process_live.py transformation ensures all amounts are divided by 10⁶ to represent standard decimal units rather than raw blockchain integers.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →