# What Data Is Included in markets.csv? Complete Schema Guide for warproxxx/poly_data

> Explore the markets.csv schema in warproxxx/poly_data. Discover 13 columns of market metadata including IDs, outcomes, volumes, and temporal data for comprehensive insights.

- Repository: [warproxxx/poly_data](https://github.com/warproxxx/poly_data)
- Tags: schema-guide
- Published: 2026-04-21

---

**The `markets.csv` file in the warproxxx/poly_data repository contains 13 columns of comprehensive metadata for every Polymarket market captured by the pipeline, including unique identifiers, outcome tokens, volume statistics, and temporal data.**

The `warproxxx/poly_data` project provides a complete data pipeline for capturing and analyzing Polymarket prediction market data. The `markets.csv` file serves as the central dataset that stores essential market metadata, making it the primary source for researchers analyzing prediction market trends and outcomes.

## Column Schema and Data Types

Each row in `markets.csv` represents a single Polymarket market with the following **13 columns** defined in [`update_utils/update_markets.py`](https://github.com/warproxxx/poly_data/blob/main/update_utils/update_markets.py) lines 33–38【/cache/repos/github.com/warproxxx/poly_data/main/update_utils/update_markets.py#L33-L38】:

- **`createdAt`** – ISO-8601 timestamp indicating when the market was created.
- **`id`** – Unique Polymarket market identifier (UUID format).
- **`question`** – Full market question text (or title if the question field is empty).
- **`answer1`** – Text of the first possible outcome (typically "YES").
- **`answer2`** – Text of the second possible outcome (typically "NO").
- **`neg_risk`** – Boolean flag indicating a negative-risk market (`True` if `negRiskAugmented` or `negRiskOther` is set).
- **`market_slug`** – URL-friendly slug for the market (e.g., `who-will-win-2024-us-presidential-election`).
- **`token1`** – CLOB token ID for the first outcome stored as a string to preserve the full 76-digit value.
- **`token2`** – CLOB token ID for the second outcome stored as a string.
- **`condition_id`** – Identifier for conditional markets (empty for standard binary markets).
- **`volume`** – Total trading volume for the market as reported by the API.
- **`ticker`** – Market ticker symbol extracted from the first event if present.
- **`closedTime`** – Timestamp when the market closed (empty if the market remains open).

The actual row assembly occurs in the same file at lines 31–45【/cache/repos/github.com/warproxxx/poly_data/main/update_utils/update_markets.py#L31-L45】, where the script populates each field from the Polymarket API response.

## How markets.csv Is Generated

The pipeline generates `markets.csv` through the **[`update_markets.py`](https://github.com/warproxxx/poly_data/blob/main/update_markets.py)** script located in the `update_utils/` directory. This script queries the Polymarket API and constructs each row using the exact column order specified in the header list.

According to the repository's **README.md** (lines 81–89)【/cache/repos/github.com/warproxxx/poly_data/main/README.md#L81-L89】, this file serves as the primary dataset combining all captured market metadata. The generation process explicitly converts numeric token IDs to strings to prevent data loss from floating-point precision issues, ensuring the full 76-digit CLOB token identifiers remain intact.

## Loading and Querying the Data

The repository provides a convenience function **`get_markets`** in [`poly_utils/utils.py`](https://github.com/warproxxx/poly_data/blob/main/poly_utils/utils.py) to load the dataset properly. This function handles schema enforcement and deduplication automatically.

### Loading with the Helper Function

```python
from poly_utils.utils import get_markets

# Load the combined market data (main and any missing markets)

markets_df = get_markets("markets.csv", "missing_markets.csv")

print(markets_df.head())
print(f"Total markets loaded: {len(markets_df)}")

```

The `get_markets` function reads the CSV using Polars and explicitly forces the **token columns to strings** at lines 20–23【/cache/repos/github.com/warproxxx/poly_data/main/poly_utils/utils.py#L20-L23】, deduplicates records based on the `id` column, and returns a DataFrame sorted by `createdAt`.

### Inspecting Market Metadata

```python

# Show all unique tickers in the dataset

print(markets_df.select("ticker").unique().sort("ticker"))

```

### Analyzing Market Creation Trends

```python

# Extract year from the `createdAt` timestamp for temporal analysis

markets_df = markets_df.with_columns(
    pl.col("createdAt").str.strptime(pl.Datetime, fmt="%Y-%m-%dT%H:%M:%S.%fZ").dt.year().alias("year")
)

print(markets_df.groupby("year").agg(pl.count()).sort("year"))

```

These patterns demonstrate standard workflows for working with the dataset once generated by the pipeline.

## Summary

- **`markets.csv`** contains 13 columns of essential Polymarket metadata defined in [`update_utils/update_markets.py`](https://github.com/warproxxx/poly_data/blob/main/update_utils/update_markets.py).
- **Token identifiers** (`token1` and `token2`) are stored as strings to preserve 76-digit precision.
- **Primary key** is the `id` column (UUID), which the loader uses for deduplication.
- **Temporal coverage** includes `createdAt` and `closedTime` for lifecycle analysis.
- **Helper function** `get_markets` in [`poly_utils/utils.py`](https://github.com/warproxxx/poly_data/blob/main/poly_utils/utils.py) provides schema-aware loading with Polars.

## Frequently Asked Questions

### What is the primary key for markets.csv?

The **`id`** column serves as the primary key, containing unique UUID values assigned by Polymarket. The `get_markets` function explicitly deduplicates on this column when loading data from multiple CSV files.

### Why are token columns stored as strings instead of integers?

The **`token1`** and **`token2`** columns contain 76-digit CLOB token identifiers that exceed standard integer precision limits. Storing them as strings prevents floating-point rounding errors and preserves the complete identifier required for blockchain interactions.

### How does the pipeline handle negative-risk markets?

The **`neg_risk`** boolean column indicates whether a market operates under negative-risk conditions. According to the source code in [`update_utils/update_markets.py`](https://github.com/warproxxx/poly_data/blob/main/update_utils/update_markets.py), this flag is set to `True` when either `negRiskAugmented` or `negRiskOther` attributes are present in the API response.

### Can I load markets.csv without using the provided helper function?

Yes, you can load the file directly with Pandas or Polars, but you must explicitly cast the **token columns to string dtype** to avoid precision loss. The `condition_id` field may be empty for standard binary markets, so handle null values appropriately in your analysis.