Fields in processed/trades.csv: Complete Schema Reference for Polymarket Trade Data

The processed/trades.csv file contains 12 standardized columns including timestamp, market_id, maker/taker addresses, directional trade flags, price, token amounts, and transaction hash.

This guide documents the complete field structure of processed/trades.csv, the primary output of the Process Live Trades pipeline in the warproxxx/poly_data repository. Whether you're analyzing Polymarket trading activity or building downstream analytics, understanding these fields is essential for correct data interpretation.


Complete Field Reference

Field Type Description
timestamp datetime Trade timestamp converted from Unix epoch
market_id string Polymarket market identifier (links to market metadata)
maker address Wallet address of the order placer
taker address Wallet address of the counterparty
nonusdc_side string Outcome token being traded (not USDC)
maker_direction "BUY" | "SELL" Direction from maker's perspective
taker_direction "BUY" | "SELL" Direction from taker's perspective
price decimal USDC per outcome token
usd_amount decimal USDC value normalized to decimal
token_amount decimal Token quantity normalized to decimal
transactionHash hex Unique blockchain transaction identifier

Where Fields Are Defined

The schema is explicitly constructed in [update_utils/process_live.py](https://github.com/warproxxx/poly_data/blob/main/update_utils/process_live.py). Lines 98-100 assemble the final DataFrame with this exact column ordering:

df = df[[
    'timestamp',
    'market_id',
    'maker',
    'taker',
    'nonusdc_side',
    'maker_direction',
    'taker_direction',
    'price',
    'usd_amount',
    'token_amount',
    'transactionHash'
]]

The same structure is documented in the repository's [README.md](https://github.com/warproxxx/poly_data/blob/main/README.md#processedtradescsv) under the processed/trades.csv section.


Working with the Data

import polars as pl

trades = (
    pl.scan_csv("processed/trades.csv")
      .collect(streaming=True)
)

print(trades.head())

Filter by market and calculate average price

market_id = "0x1234567890abcdef"

avg_price = (
    trades
    .filter(pl.col("market_id") == market_id)
    .select(pl.col("price").mean())
    .item()
)

print(f"Average price: {avg_price:.4f} USDC per token")

Aggregate volume by trader

maker_volumes = (
    trades
    .groupby("maker")
    .agg(pl.col("usd_amount").sum().alias("total_usd"))
    .sort("total_usd", descending=True)
)

print(maker_volumes.head())

Key Design Decisions

Understanding why certain fields exist helps prevent analysis errors:

  • nonusdc_side — Polymarket trades always involve USDC on one side. This field identifies which outcome token is actually changing hands, simplifying position calculations.

  • Directional duality (maker_direction/taker_direction) — These are inverses of each other. The maker's buy is always the taker's sell. Both are included because different analyses care about different perspectives.

  • Decimal normalization — usd_amount and token_amount are converted from raw blockchain integers to human-readable decimals using each token's specific decimals value, eliminating manual conversion errors.


Summary

  • processed/trades.csv in the warproxxx/poly_data repository contains 12 standardized fields for Polymarket trade analysis
  • Fields are explicitly defined in update_utils/process_live.py lines 98-100 and documented in README.md
  • Key fields include market_id, maker/taker addresses, maker_direction/taker_direction, price, and transactionHash
  • nonusdc_side identifies the traded outcome token since Polymarket always involves USDC on one side
  • Use Polars with streaming=True for efficient processing of large trade files

Frequently Asked Questions

What is the difference between maker_direction and taker_direction?

These fields represent opposite sides of the same trade. maker_direction shows whether the maker (order placer) is buying or selling, while taker_direction shows the same from the counterparty's perspective. If the maker is buying, the taker is necessarily selling. Both are included to support analyses from either trader's viewpoint.

Why is there a nonusdc_side field instead of just listing both tokens?

Polymarket's design fixes USDC as one side of every trade. The nonusdc_side field identifies which outcome token is actually being exchanged, eliminating redundancy and making position calculations more intuitive. This design choice reflects the protocol's token structure where USDC serves as the universal quote currency.

How are usd_amount and token_amount calculated from raw blockchain data?

These fields are normalized from raw integer values using each token's specific decimal precision. The pipeline in update_utils/process_live.py applies the appropriate 10**decimals conversion factor, producing human-readable values. This eliminates manual conversion errors and ensures consistent scaling across different tokens with varying decimal places.

Where can I find the exact column order for processed/trades.csv?

The column ordering is explicitly defined in update_utils/process_live.py at lines 98-100, where the final DataFrame is assembled with a hardcoded list of 12 field names. The same schema is documented in the repository's README.md under the processed/trades.csv section. Reference these sources when building dependent systems that expect specific column positions.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →