What Data Is Included in markets.csv? Complete Schema Guide for warproxxx/poly_data

The markets.csv file in the warproxxx/poly_data repository contains 13 columns of comprehensive metadata for every Polymarket market captured by the pipeline, including unique identifiers, outcome tokens, volume statistics, and temporal data.

The warproxxx/poly_data project provides a complete data pipeline for capturing and analyzing Polymarket prediction market data. The markets.csv file serves as the central dataset that stores essential market metadata, making it the primary source for researchers analyzing prediction market trends and outcomes.

Column Schema and Data Types

Each row in markets.csv represents a single Polymarket market with the following 13 columns defined in update_utils/update_markets.py lines 33–38【/cache/repos/github.com/warproxxx/poly_data/main/update_utils/update_markets.py#L33-L38】:

  • createdAt – ISO-8601 timestamp indicating when the market was created.
  • id – Unique Polymarket market identifier (UUID format).
  • question – Full market question text (or title if the question field is empty).
  • answer1 – Text of the first possible outcome (typically "YES").
  • answer2 – Text of the second possible outcome (typically "NO").
  • neg_risk – Boolean flag indicating a negative-risk market (True if negRiskAugmented or negRiskOther is set).
  • market_slug – URL-friendly slug for the market (e.g., who-will-win-2024-us-presidential-election).
  • token1 – CLOB token ID for the first outcome stored as a string to preserve the full 76-digit value.
  • token2 – CLOB token ID for the second outcome stored as a string.
  • condition_id – Identifier for conditional markets (empty for standard binary markets).
  • volume – Total trading volume for the market as reported by the API.
  • ticker – Market ticker symbol extracted from the first event if present.
  • closedTime – Timestamp when the market closed (empty if the market remains open).

The actual row assembly occurs in the same file at lines 31–45【/cache/repos/github.com/warproxxx/poly_data/main/update_utils/update_markets.py#L31-L45】, where the script populates each field from the Polymarket API response.

How markets.csv Is Generated

The pipeline generates markets.csv through the update_markets.py script located in the update_utils/ directory. This script queries the Polymarket API and constructs each row using the exact column order specified in the header list.

According to the repository's README.md (lines 81–89)【/cache/repos/github.com/warproxxx/poly_data/main/README.md#L81-L89】, this file serves as the primary dataset combining all captured market metadata. The generation process explicitly converts numeric token IDs to strings to prevent data loss from floating-point precision issues, ensuring the full 76-digit CLOB token identifiers remain intact.

Loading and Querying the Data

The repository provides a convenience function get_markets in poly_utils/utils.py to load the dataset properly. This function handles schema enforcement and deduplication automatically.

Loading with the Helper Function

from poly_utils.utils import get_markets

# Load the combined market data (main and any missing markets)

markets_df = get_markets("markets.csv", "missing_markets.csv")

print(markets_df.head())
print(f"Total markets loaded: {len(markets_df)}")

The get_markets function reads the CSV using Polars and explicitly forces the token columns to strings at lines 20–23【/cache/repos/github.com/warproxxx/poly_data/main/poly_utils/utils.py#L20-L23】, deduplicates records based on the id column, and returns a DataFrame sorted by createdAt.

Inspecting Market Metadata


# Show all unique tickers in the dataset

print(markets_df.select("ticker").unique().sort("ticker"))

# Extract year from the `createdAt` timestamp for temporal analysis

markets_df = markets_df.with_columns(
    pl.col("createdAt").str.strptime(pl.Datetime, fmt="%Y-%m-%dT%H:%M:%S.%fZ").dt.year().alias("year")
)

print(markets_df.groupby("year").agg(pl.count()).sort("year"))

These patterns demonstrate standard workflows for working with the dataset once generated by the pipeline.

Summary

  • markets.csv contains 13 columns of essential Polymarket metadata defined in update_utils/update_markets.py.
  • Token identifiers (token1 and token2) are stored as strings to preserve 76-digit precision.
  • Primary key is the id column (UUID), which the loader uses for deduplication.
  • Temporal coverage includes createdAt and closedTime for lifecycle analysis.
  • Helper function get_markets in poly_utils/utils.py provides schema-aware loading with Polars.

Frequently Asked Questions

What is the primary key for markets.csv?

The id column serves as the primary key, containing unique UUID values assigned by Polymarket. The get_markets function explicitly deduplicates on this column when loading data from multiple CSV files.

Why are token columns stored as strings instead of integers?

The token1 and token2 columns contain 76-digit CLOB token identifiers that exceed standard integer precision limits. Storing them as strings prevents floating-point rounding errors and preserves the complete identifier required for blockchain interactions.

How does the pipeline handle negative-risk markets?

The neg_risk boolean column indicates whether a market operates under negative-risk conditions. According to the source code in update_utils/update_markets.py, this flag is set to True when either negRiskAugmented or negRiskOther attributes are present in the API response.

Can I load markets.csv without using the provided helper function?

Yes, you can load the file directly with Pandas or Polars, but you must explicitly cast the token columns to string dtype to avoid precision loss. The condition_id field may be empty for standard binary markets, so handle null values appropriately in your analysis.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →