# How Asset IDs Are Mapped to Markets in Poly Data

> Discover how Poly Data maps asset IDs to markets using token identifiers in CSV files and a long lookup table for efficient market data retrieval. Learn the process now.

- Repository: [warproxxx/poly_data](https://github.com/warproxxx/poly_data)
- Tags: internals
- Published: 2026-04-21

---

**Poly Data maps asset IDs to markets by storing `token1` and `token2` identifiers in market CSV files, then reshaping this data into a long lookup table that joins each non‑USDC asset ID to its corresponding `market_id` and side.**

The Poly Data library, maintained at `warproxxx/poly_data`, provides a robust system for resolving which market any given asset belongs to. This mapping is essential for processing live trade data from the Polymarket ecosystem, where every trade involves two assets but only one represents the actual market token (the other is typically USDC, represented as `"0"`).

## The Core Mapping Mechanism

The foundation of asset-to-market mapping lies in two simple columns: **`token1`** and **`token2`**. These columns exist in every market record and store the unique identifier for each asset traded in that market.

When processing trades, the system cannot directly use the wide format (where both tokens sit in the same row). Instead, it transforms this structure into a **long format** where each asset ID occupies its own row, paired with its `market_id` and a `side` indicator showing whether it is `token1` or `token2`.

## How the Mapping Works in Practice

The complete workflow is implemented in [`update_utils/process_live.py`](https://github.com/warproxxx/poly_data/blob/main/update_utils/process_live.py). Here is the step-by-step process:

### 1. Load Market Data

The `get_markets()` function from [`poly_utils/utils.py`](https://github.com/warproxxx/poly_data/blob/main/poly_utils/utils.py) reads the combined market files and returns a Polars DataFrame containing `market_id`, `token1`, and `token2`:

```python
markets = get_markets().rename({"id": "market_id"})

```

### 2. Reshape to Long Form

The `melt()` operation transforms the wide structure into rows of `(market_id, side, asset_id)`:

```python
markets_long = (
    markets.select(["market_id", "token1", "token2"])
            .melt(id_vars="market_id",
                  value_vars=["token1", "token2"],
                  variable_name="side",
                  value_name="asset_id")
)

```

### 3. Identify the Non-USDC Asset

In each trade, one asset is USDC (ID `"0"`) and the other is the actual market token. The code creates a `nonusdc_asset_id` column:

```python
trades_df = trades_df.with_columns(
    pl.when(pl.col("makerAssetId") != "0")
      .then(pl.col("makerAssetId"))
      .otherwise(pl.col("takerAssetId"))
      .alias("nonusdc_asset_id")
)

```

### 4. Join to Resolve Market ID

The final join connects each trade to its market by matching `nonusdc_asset_id` against `asset_id` in the long lookup table:

```python
enriched = trades_df.join(
    markets_long,
    left_on="nonusdc_asset_id",
    right_on="asset_id",
    how="left"
)

```

After this join, every trade row contains `market_id` and `side`, fully resolving which market each asset belongs to.

## Complete Working Example

Here is a reproducible function that implements the entire mapping pipeline:

```python
import polars as pl
from poly_utils.utils import get_markets

def map_asset_ids_to_markets(trades_df: pl.DataFrame) -> pl.DataFrame:
    """
    Maps asset IDs in trade records to their corresponding market IDs.
    
    Parameters
    ----------
    trades_df : pl.DataFrame
        Raw trades with makerAssetId and takerAssetId columns
        
    Returns
    -------
    pl.DataFrame
        Enriched trades with market_id and side columns
    """
    # Build the long-form asset-to-market lookup table

    markets = get_markets().rename({"id": "market_id"})
    markets_long = (
        markets.select(["market_id", "token1", "token2"])
                .melt(
                    id_vars="market_id",
                    value_vars=["token1", "token2"],
                    variable_name="side",
                    value_name="asset_id"
                )
    )
    
    # Identify which asset in each trade is the market token (non-USDC)

    trades_with_asset = trades_df.with_columns(
        pl.when(pl.col("makerAssetId") != "0")
          .then(pl.col("makerAssetId"))
          .otherwise(pl.col("takerAssetId"))
          .alias("nonusdc_asset_id")
    )
    
    # Join to resolve market_id and side

    result = trades_with_asset.join(
        markets_long,
        left_on="nonusdc_asset_id",
        right_on="asset_id",
        how="left"
    )
    
    return result

```

## Key Source Files

| File | Purpose |
|------|---------|
| [`update_utils/process_live.py`](https://github.com/warproxxx/poly_data/blob/main/update_utils/process_live.py) | Implements the live-trade processing pipeline and performs the asset-to-market join |
| [`poly_utils/utils.py`](https://github.com/warproxxx/poly_data/blob/main/poly_utils/utils.py) | Provides `get_markets()` which loads market data and supplies the `token1` / `token2` asset columns |

## Summary

- **Asset IDs map to markets through `token1` and `token2` columns** in market CSV files
- **The long-form lookup table** created by `melt()` enables efficient joins from any asset ID to its market
- **Non-USDC asset identification** resolves which side of a trade contains the actual market token
- **A single Polars join** in [`update_utils/process_live.py`](https://github.com/warproxxx/poly_data/blob/main/update_utils/process_live.py) completes the mapping, adding `market_id` and `side` to each trade record

## Frequently Asked Questions

### What are `token1` and `token2` in Poly Data?

`token1` and `token2` are column names in the market CSV files that store the unique asset identifiers for the two tokens traded in each market. One of these is typically USDC (represented as `"0"`), while the other represents the actual prediction market token.

### How does Poly Data handle trades where both assets are non-USDC?

The current implementation in [`update_utils/process_live.py`](https://github.com/warproxxx/poly_data/blob/main/update_utils/process_live.py) uses a simple conditional: it checks if `makerAssetId != "0"` and uses that value; otherwise, it falls back to `takerAssetId`. This assumes one side is always USDC. For markets with two non-USDC assets, the logic would need modification to explicitly match both asset IDs against the market lookup table.

### Where is the market data loaded from in Poly Data?

The `get_markets()` function in [`poly_utils/utils.py`](https://github.com/warproxxx/poly_data/blob/main/poly_utils/utils.py) loads market data from `markets.csv` and `missing_markets.csv` files. These CSVs contain the `market_id`, `token1`, and `token2` columns that enable the asset-to-market mapping throughout the codebase.

### Why use `melt()` instead of keeping the wide format?

The `melt()` operation transforms the wide format (where both tokens share a row) into a long format where each asset ID has its own row with `market_id` and `side`. This structure is essential for performing efficient joins: it allows any single asset ID from a trade to match directly against the lookup table without complex conditional logic or multiple join attempts.