# How Poly Data Identifies USDC in Trades: A Technical Deep Dive

> Learn how Poly Data identifies USDC in trades by treating asset ID 0 as USDC and using Polars conditional expressions for accurate labeling.

- Repository: [warproxxx/poly_data](https://github.com/warproxxx/poly_data)
- Tags: deep-dive
- Published: 2026-04-21

---

**Poly Data identifies USDC in trades by treating asset ID `"0"` as the USDC token, then using Polars conditional expressions to label maker and taker assets accordingly.**

The Poly Data repository processes raw decentralized exchange data—primarily from the 0x protocol—to produce clean, analysis-ready datasets. A critical step in this pipeline is correctly identifying which side of a trade involves USDC, the dominant stablecoin in most DeFi markets. This article explains exactly how Poly Data identifies USDC in trades, walking through the source code in [`update_utils/process_live.py`](https://github.com/warproxxx/poly_data/blob/main/update_utils/process_live.py).

## The Core Convention: Asset ID "0" Equals USDC

Poly Data relies on a simple but strict convention: **the string `"0"` represents USDC** in all raw trade data. This design choice eliminates the need for external token contract lookups during processing and ensures consistent labeling across all markets.

When either `makerAssetId` or `takerAssetId` equals `"0"`, Poly Data immediately flags that side as the USDC leg of the trade. This detection happens inside the `get_processed_df` function, which transforms raw Polars DataFrames into enriched, labeled datasets.

## Step-by-Step USDC Identification in [`process_live.py`](https://github.com/warproxxx/poly_data/blob/main/process_live.py)

The USDC identification logic follows a five-step pipeline implemented in [`update_utils/process_live.py`](https://github.com/warproxxx/poly_data/blob/main/update_utils/process_live.py). Each step builds on the previous to create explicit, queryable columns.

### Step 1: Isolate the Non-USDC Asset

First, Poly Data creates a helper column `nonusdc_asset_id` that stores whichever asset ID is **not** `"0"`. This simplifies downstream joins and calculations.

```python

# Conceptual logic—actual implementation uses Polars expressions

df = df.with_columns([
    pl.when(pl.col("makerAssetId") != "0")
      .then(pl.col("makerAssetId"))
      .otherwise(pl.col("takerAssetId"))
      .alias("nonusdc_asset_id")
])

```

### Step 2: Join with Market Definitions

Next, the pipeline joins trade rows with the market-definition table via [`poly_utils/utils.py`](https://github.com/warproxxx/poly_data/blob/main/poly_utils/utils.py)'s `get_markets()` function. This retrieval step maps the `nonusdc_asset_id` to its `market_id` and determines whether it represents `token1` or `token2` in that market's structure.

### Step 3: Label Maker and Taker Assets

The core USDC identification happens here. Using Polars' `pl.when().then().otherwise()` conditional expressions, Poly Data explicitly labels each side:

```python

# Lines 45-46 of update_utils/process_live.py (excerpt)

df = df.with_columns([
    # Maker asset: "USDC" if ID is "0", otherwise the side name

    pl.when(pl.col("makerAssetId") == "0")
      .then(pl.lit("USDC"))
      .otherwise(pl.col("side"))
      .alias("makerAsset"),

    # Taker asset: same logic

    pl.when(pl.col("takerAssetId") == "0")
      .then(pl.lit("USDC"))
      .otherwise(pl.col("side"))
      .alias("takerAsset")
])

```

This produces human-readable columns where `"USDC"` appears explicitly when the asset ID was `"0"`, and `"token1"` or `"token2"` appears otherwise.

### Step 4: Derive Trade Direction

With USDC identified, Poly Data computes trade direction. The side receiving USDC is considered a **BUY** for the taker (and **SELL** for the maker), with the inverse when USDC is on the maker side. Additional conditional columns `taker_direction` and `maker_direction` capture this.

### Step 5: Create Convenience Columns

Finally, the pipeline exposes clean, analysis-ready columns:

- `nonusdc_side`: which side (`token1`/`token2`) represents the non-USDC asset
- `usd_amount`: the USDC-denominated value of the trade
- `token_amount`: the quantity of the non-USDC asset
- `price`: derived as `usd_amount / token_amount`

These columns eliminate the need to reference raw asset IDs for downstream analytics.

## Complete Processing Example

Here's how to use the complete pipeline:

```python
from update_utils.process_live import get_processed_df
import polars as pl

# Raw trade data from 0x protocol events

raw_df = pl.DataFrame({
    "makerAssetId": ["0", "12345", "0"],
    "takerAssetId": ["67890", "0", "54321"],
    "makerAmount": ["1000000000", "500000000", "250000000"],
    "takerAmount": ["1500000000000", "750000000000", "500000000000"]
})

# Process with USDC identification built in

clean_df = get_processed_df(raw_df)

# Result includes explicit USDC labeling

print(clean_df.select(["makerAsset", "takerAsset", "maker_direction", "taker_direction"]))

```

## Key Implementation Files

| File | Purpose |
|------|---------|
| [`update_utils/process_live.py`](https://github.com/warproxxx/poly_data/blob/main/update_utils/process_live.py) | Core trade processing with USDC identification via `get_processed_df()` |
| [`poly_utils/utils.py`](https://github.com/warproxxx/poly_data/blob/main/poly_utils/utils.py) | Market definitions via `get_markets()` for asset-to-side mapping |
| [`README.md`](https://github.com/warproxxx/poly_data/blob/main/README.md) | Pipeline overview and repository documentation |

## Summary

- **Poly Data identifies USDC by treating asset ID `"0"` as the USDC token** in all raw trade data.
- The `get_processed_df()` function in [`update_utils/process_live.py`](https://github.com/warproxxx/poly_data/blob/main/update_utils/process_live.py) implements a five-step pipeline: isolate non-USDC assets, join market definitions, label maker/taker assets with Polars conditionals, derive trade direction, and create convenience columns.
- **Explicit `"USDC"` strings** appear in output columns via `pl.when(col == "0").then("USDC").otherwise(side)`, making downstream analysis straightforward.

## Frequently Asked Questions

### How does Poly Data handle trades where neither asset is USDC?

Poly Data's current implementation assumes at least one side of every trade is USDC. If neither `makerAssetId` nor `takerAssetId` equals `"0"`, both `makerAsset` and `takerAsset` would receive the respective `side` values (`token1` or `token2`), and USDC-specific columns like `usd_amount` would not populate correctly. This design reflects the repository's focus on USDC-quoted markets.

### Can the USDC asset ID be configured to something other than "0"?

The asset ID `"0"` is **hardcoded** throughout the pipeline. In [`update_utils/process_live.py`](https://github.com/warproxxx/poly_data/blob/main/update_utils/process_live.py), the conditional expressions explicitly check `pl.col("makerAssetId") == "0"` and `pl.col("takerAssetId") == "0"`. Changing the USDC identifier would require modifying these expressions and any related logic in [`poly_utils/utils.py`](https://github.com/warproxxx/poly_data/blob/main/poly_utils/utils.py) that assumes the `"0"` convention.

### What Polars operations enable the conditional USDC labeling?

Poly Data uses **Polars' `when-then-otherwise` expressions** for all conditional logic. The pattern `pl.when(condition).then(value_if_true).otherwise(value_if_false)` appears throughout `get_processed_df()`, most critically for setting `makerAsset` and `takerAsset` to `"USDC"` when the respective asset ID equals `"0"`. This approach vectorizes operations across the entire DataFrame without Python loops.