# How Are Amounts Normalized in Poly Data Trades: A Complete Technical Guide

> Learn how Poly Data normalizes trade amounts by dividing raw on-chain integer values by 10⁶ in update_utils/process_live.py converting them to human-readable decimals.

- Repository: [warproxxx/poly_data](https://github.com/warproxxx/poly_data)
- Tags: deep-dive
- Published: 2026-04-21

---

**Poly Data normalizes trade amounts by dividing raw on-chain integer values by 10⁶ inside the `get_processed_df` function in [`update_utils/process_live.py`](https://github.com/warproxxx/poly_data/blob/main/update_utils/process_live.py), converting token quantities from their smallest unit representation into human-readable decimal format.**

The `warproxxx/poly_data` repository processes decentralized exchange trade data from Goldsky's `orderFilled.csv` exports. Understanding how amounts are normalized in Poly Data trades is essential for anyone analyzing the processed output, as the conversion determines the accuracy of subsequent price and volume calculations.

## The Normalization Logic in [`process_live.py`](https://github.com/warproxxx/poly_data/blob/main/process_live.py)

All raw amount normalization occurs within **[`update_utils/process_live.py`](https://github.com/warproxxx/poly_data/blob/main/update_utils/process_live.py)** in the **`get_processed_df`** function. The code assumes a standard **6 decimal place** precision (typical for USDC and most ERC-20 tokens) and applies this transformation using Polars DataFrame operations.

The core normalization divides both maker and taker amounts by 10⁶:

```python
df = df.with_columns([
    (pl.col("makerAmountFilled") / 10**6).alias("makerAmountFilled"),
    (pl.col("takerAmountFilled") / 10**6).alias("takerAmountFilled"),
])

```

This operation transforms integer values such as `1500000` into `1.5`, representing the actual token quantity readable by humans. The division happens immediately after loading the raw Goldsky data, ensuring all downstream calculations use decimal-adjusted values.

## Building Derived Metrics from Normalized Amounts

After the initial division by 10⁶, the script constructs three critical derived columns using conditional logic based on which side of the trade involves USDC:

- **`usd_amount`**: Captures the USDC-denominated side of the trade
- **`token_amount`**: Captures the non-USDC token quantity  
- **`price`**: Calculates the exchange rate between the two assets

The implementation checks the `takerAsset` column to identify the USDC side:

```python
df = df.with_columns([
    pl.when(pl.col("takerAsset") == "USDC")
      .then(pl.col("takerAmountFilled"))
      .otherwise(pl.col("makerAmountFilled"))
      .alias("usd_amount"),

    pl.when(pl.col("takerAsset") != "USDC")
      .then(pl.col("takerAmountFilled"))
      .otherwise(pl.col("makerAmountFilled"))
      .alias("token_amount"),

    pl.when(pl.col("takerAsset") == "USDC")
      .then(pl.col("takerAmountFilled") / pl.col("makerAmountFilled"))
      .otherwise(pl.col("makerAmountFilled") / pl.col("takerAmountFilled"))
      .cast(pl.Float64)
      .alias("price"),
])

```

The final DataFrame selects these normalized and derived columns for output to `processed/trades.csv`:

```python
df = df[['timestamp', 'market_id', 'maker', 'taker',
        'nonusdc_side', 'maker_direction', 'taker_direction',
        'price', 'usd_amount', 'token_amount', 'transactionHash']]

```

## Input Data and Decimal Assumptions

The normalization logic relies on data exported from **Goldsky** containing on-chain event logs. The code explicitly assumes all tokens utilize **6 decimal places**, which aligns with USDC and many stablecoin implementations. This assumption is hardcoded through the literal `10**6` divisor rather than being dynamically fetched from token metadata.

If analyzing tokens with different decimal precisions (such as 18-decimal WETH or WBTC), the current implementation would require modification to reference the specific token's `decimals` field, as the fixed 10⁶ division would otherwise miscalculate the actual transferred amounts.

## Summary

- **Primary transformation**: Raw amounts are divided by 10⁶ in [`update_utils/process_live.py`](https://github.com/warproxxx/poly_data/blob/main/update_utils/process_live.py) to convert from integer smallest-unit representation to decimal values.
- **Key function**: `get_processed_df` handles both the normalization and the creation of derived metrics.
- **Derived columns**: `usd_amount`, `token_amount`, and `price` are calculated after normalization using conditional logic based on which asset is USDC.
- **Data source**: Normalization applies to Goldsky `orderFilled.csv` exports containing raw on-chain trade events.
- **Assumption**: The code assumes 6 decimal places for all tokens, matching USDC and standard ERC-20 stablecoin precision.

## Frequently Asked Questions

### Why does Poly Data divide amounts by 10⁶ specifically?

The division by 10⁶ converts raw blockchain integers into decimal representation based on the assumption that all tokens use 6 decimal places. This matches USDC and many stablecoins where `1000000` on-chain units equal `1.0` actual tokens. The value is hardcoded in [`process_live.py`](https://github.com/warproxxx/poly_data/blob/main/process_live.py) because the pipeline primarily processes USDC-based trading pairs.

### What happens if a token has 18 decimals instead of 6?

The current normalization would produce incorrect values for 18-decimal tokens, displaying amounts that are 10¹² times smaller than actual. For example, `1000000000000000000` (1 WETH) would appear as `1,000,000` instead of `1.0`. The repository currently assumes uniform 6-decimal precision across all traded assets.

### Where does the raw trade data originate before normalization?

Raw data comes from **Goldsky** CSV exports specifically from the `orderFilled.csv` file, which contains on-chain event logs from the Poly Market smart contracts. These files provide integer values for `makerAmountFilled` and `takerAmountFilled` that require decimal adjustment before analysis.

### How is the trade price calculated after amount normalization?

The price is computed as a ratio of the normalized amounts after the division by 10⁶. The code uses Polars `when/then/otherwise` logic to divide the USDC amount by the token amount (or vice versa depending on trade direction), casting the result to `Float64` for precision. This calculation depends entirely on the pre-normalized values being correctly scaled to their decimal representations.