How Asset IDs Are Mapped to Markets in Poly Data
Poly Data maps asset IDs to markets by storing token1 and token2 identifiers in market CSV files, then reshaping this data into a long lookup table that joins each non‑USDC asset ID to its corresponding market_id and side.
The Poly Data library, maintained at warproxxx/poly_data, provides a robust system for resolving which market any given asset belongs to. This mapping is essential for processing live trade data from the Polymarket ecosystem, where every trade involves two assets but only one represents the actual market token (the other is typically USDC, represented as "0").
The Core Mapping Mechanism
The foundation of asset-to-market mapping lies in two simple columns: token1 and token2. These columns exist in every market record and store the unique identifier for each asset traded in that market.
When processing trades, the system cannot directly use the wide format (where both tokens sit in the same row). Instead, it transforms this structure into a long format where each asset ID occupies its own row, paired with its market_id and a side indicator showing whether it is token1 or token2.
How the Mapping Works in Practice
The complete workflow is implemented in update_utils/process_live.py. Here is the step-by-step process:
1. Load Market Data
The get_markets() function from poly_utils/utils.py reads the combined market files and returns a Polars DataFrame containing market_id, token1, and token2:
markets = get_markets().rename({"id": "market_id"})
2. Reshape to Long Form
The melt() operation transforms the wide structure into rows of (market_id, side, asset_id):
markets_long = (
markets.select(["market_id", "token1", "token2"])
.melt(id_vars="market_id",
value_vars=["token1", "token2"],
variable_name="side",
value_name="asset_id")
)
3. Identify the Non-USDC Asset
In each trade, one asset is USDC (ID "0") and the other is the actual market token. The code creates a nonusdc_asset_id column:
trades_df = trades_df.with_columns(
pl.when(pl.col("makerAssetId") != "0")
.then(pl.col("makerAssetId"))
.otherwise(pl.col("takerAssetId"))
.alias("nonusdc_asset_id")
)
4. Join to Resolve Market ID
The final join connects each trade to its market by matching nonusdc_asset_id against asset_id in the long lookup table:
enriched = trades_df.join(
markets_long,
left_on="nonusdc_asset_id",
right_on="asset_id",
how="left"
)
After this join, every trade row contains market_id and side, fully resolving which market each asset belongs to.
Complete Working Example
Here is a reproducible function that implements the entire mapping pipeline:
import polars as pl
from poly_utils.utils import get_markets
def map_asset_ids_to_markets(trades_df: pl.DataFrame) -> pl.DataFrame:
"""
Maps asset IDs in trade records to their corresponding market IDs.
Parameters
----------
trades_df : pl.DataFrame
Raw trades with makerAssetId and takerAssetId columns
Returns
-------
pl.DataFrame
Enriched trades with market_id and side columns
"""
# Build the long-form asset-to-market lookup table
markets = get_markets().rename({"id": "market_id"})
markets_long = (
markets.select(["market_id", "token1", "token2"])
.melt(
id_vars="market_id",
value_vars=["token1", "token2"],
variable_name="side",
value_name="asset_id"
)
)
# Identify which asset in each trade is the market token (non-USDC)
trades_with_asset = trades_df.with_columns(
pl.when(pl.col("makerAssetId") != "0")
.then(pl.col("makerAssetId"))
.otherwise(pl.col("takerAssetId"))
.alias("nonusdc_asset_id")
)
# Join to resolve market_id and side
result = trades_with_asset.join(
markets_long,
left_on="nonusdc_asset_id",
right_on="asset_id",
how="left"
)
return result
Key Source Files
| File | Purpose |
|---|---|
update_utils/process_live.py |
Implements the live-trade processing pipeline and performs the asset-to-market join |
poly_utils/utils.py |
Provides get_markets() which loads market data and supplies the token1 / token2 asset columns |
Summary
- Asset IDs map to markets through
token1andtoken2columns in market CSV files - The long-form lookup table created by
melt()enables efficient joins from any asset ID to its market - Non-USDC asset identification resolves which side of a trade contains the actual market token
- A single Polars join in
update_utils/process_live.pycompletes the mapping, addingmarket_idandsideto each trade record
Frequently Asked Questions
What are token1 and token2 in Poly Data?
token1 and token2 are column names in the market CSV files that store the unique asset identifiers for the two tokens traded in each market. One of these is typically USDC (represented as "0"), while the other represents the actual prediction market token.
How does Poly Data handle trades where both assets are non-USDC?
The current implementation in update_utils/process_live.py uses a simple conditional: it checks if makerAssetId != "0" and uses that value; otherwise, it falls back to takerAssetId. This assumes one side is always USDC. For markets with two non-USDC assets, the logic would need modification to explicitly match both asset IDs against the market lookup table.
Where is the market data loaded from in Poly Data?
The get_markets() function in poly_utils/utils.py loads market data from markets.csv and missing_markets.csv files. These CSVs contain the market_id, token1, and token2 columns that enable the asset-to-market mapping throughout the codebase.
Why use melt() instead of keeping the wide format?
The melt() operation transforms the wide format (where both tokens share a row) into a long format where each asset ID has its own row with market_id and side. This structure is essential for performing efficient joins: it allows any single asset ID from a trade to match directly against the lookup table without complex conditional logic or multiple join attempts.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →