How Are Amounts Normalized in Poly Data Trades: A Complete Technical Guide
Poly Data normalizes trade amounts by dividing raw on-chain integer values by 10⁶ inside the get_processed_df function in update_utils/process_live.py, converting token quantities from their smallest unit representation into human-readable decimal format.
The warproxxx/poly_data repository processes decentralized exchange trade data from Goldsky's orderFilled.csv exports. Understanding how amounts are normalized in Poly Data trades is essential for anyone analyzing the processed output, as the conversion determines the accuracy of subsequent price and volume calculations.
The Normalization Logic in process_live.py
All raw amount normalization occurs within update_utils/process_live.py in the get_processed_df function. The code assumes a standard 6 decimal place precision (typical for USDC and most ERC-20 tokens) and applies this transformation using Polars DataFrame operations.
The core normalization divides both maker and taker amounts by 10⁶:
df = df.with_columns([
(pl.col("makerAmountFilled") / 10**6).alias("makerAmountFilled"),
(pl.col("takerAmountFilled") / 10**6).alias("takerAmountFilled"),
])
This operation transforms integer values such as 1500000 into 1.5, representing the actual token quantity readable by humans. The division happens immediately after loading the raw Goldsky data, ensuring all downstream calculations use decimal-adjusted values.
Building Derived Metrics from Normalized Amounts
After the initial division by 10⁶, the script constructs three critical derived columns using conditional logic based on which side of the trade involves USDC:
usd_amount: Captures the USDC-denominated side of the tradetoken_amount: Captures the non-USDC token quantityprice: Calculates the exchange rate between the two assets
The implementation checks the takerAsset column to identify the USDC side:
df = df.with_columns([
pl.when(pl.col("takerAsset") == "USDC")
.then(pl.col("takerAmountFilled"))
.otherwise(pl.col("makerAmountFilled"))
.alias("usd_amount"),
pl.when(pl.col("takerAsset") != "USDC")
.then(pl.col("takerAmountFilled"))
.otherwise(pl.col("makerAmountFilled"))
.alias("token_amount"),
pl.when(pl.col("takerAsset") == "USDC")
.then(pl.col("takerAmountFilled") / pl.col("makerAmountFilled"))
.otherwise(pl.col("makerAmountFilled") / pl.col("takerAmountFilled"))
.cast(pl.Float64)
.alias("price"),
])
The final DataFrame selects these normalized and derived columns for output to processed/trades.csv:
df = df[['timestamp', 'market_id', 'maker', 'taker',
'nonusdc_side', 'maker_direction', 'taker_direction',
'price', 'usd_amount', 'token_amount', 'transactionHash']]
Input Data and Decimal Assumptions
The normalization logic relies on data exported from Goldsky containing on-chain event logs. The code explicitly assumes all tokens utilize 6 decimal places, which aligns with USDC and many stablecoin implementations. This assumption is hardcoded through the literal 10**6 divisor rather than being dynamically fetched from token metadata.
If analyzing tokens with different decimal precisions (such as 18-decimal WETH or WBTC), the current implementation would require modification to reference the specific token's decimals field, as the fixed 10⁶ division would otherwise miscalculate the actual transferred amounts.
Summary
- Primary transformation: Raw amounts are divided by 10⁶ in
update_utils/process_live.pyto convert from integer smallest-unit representation to decimal values. - Key function:
get_processed_dfhandles both the normalization and the creation of derived metrics. - Derived columns:
usd_amount,token_amount, andpriceare calculated after normalization using conditional logic based on which asset is USDC. - Data source: Normalization applies to Goldsky
orderFilled.csvexports containing raw on-chain trade events. - Assumption: The code assumes 6 decimal places for all tokens, matching USDC and standard ERC-20 stablecoin precision.
Frequently Asked Questions
Why does Poly Data divide amounts by 10⁶ specifically?
The division by 10⁶ converts raw blockchain integers into decimal representation based on the assumption that all tokens use 6 decimal places. This matches USDC and many stablecoins where 1000000 on-chain units equal 1.0 actual tokens. The value is hardcoded in process_live.py because the pipeline primarily processes USDC-based trading pairs.
What happens if a token has 18 decimals instead of 6?
The current normalization would produce incorrect values for 18-decimal tokens, displaying amounts that are 10¹² times smaller than actual. For example, 1000000000000000000 (1 WETH) would appear as 1,000,000 instead of 1.0. The repository currently assumes uniform 6-decimal precision across all traded assets.
Where does the raw trade data originate before normalization?
Raw data comes from Goldsky CSV exports specifically from the orderFilled.csv file, which contains on-chain event logs from the Poly Market smart contracts. These files provide integer values for makerAmountFilled and takerAmountFilled that require decimal adjustment before analysis.
How is the trade price calculated after amount normalization?
The price is computed as a ratio of the normalized amounts after the division by 10⁶. The code uses Polars when/then/otherwise logic to divide the USDC amount by the token amount (or vice versa depending on trade direction), casting the result to Float64 for precision. This calculation depends entirely on the pre-normalized values being correctly scaled to their decimal representations.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →