How Kronos Handles Missing Volume and Amount Data in Input DataFrames

Kronos automatically detects missing volume and amount columns in input DataFrames and repairs them by inserting zero placeholders or deriving amount from the mean price, ensuring a complete OHLCVA feature set before model inference.

Financial time series data often arrives incomplete. The shiyu-coder/Kronos repository is designed to tolerate such gaps by preprocessing inputs inside the KronosPredictor class. When you call predict() or predict_batch(), the library guarantees six complete numeric columns—open, high, low, close, volume, and amount—before any autoregressive computation begins.

Detecting Missing Volume and Amount Columns

Inside model/kronos.py, the KronosPredictor class inspects incoming DataFrames for the presence of optional columns mapped to self.vol_col (volume) and self.amt_vol (amount). Rather than failing immediately when these fields are absent, the predictor enters a repair pipeline that prepares the data for downstream normalization and tokenization.

Repair Strategies for Incomplete Market Data

Inserting Zero Placeholders for Absent Columns

When Kronos identifies missing volume and amount data simultaneously, it creates both columns and fills every row with 0.0. This ensures the feature matrix maintains consistent dimensions and prevents null reference errors during tensor conversion.

Deriving Amount from Available Volume

If the DataFrame contains a volume column but lacks amount, Kronos computes a proxy turnover value using the average of the price columns. According to the source implementation in model/kronos.py, the calculation follows:


amount = volume × mean([open, high, low, close])

This derivation provides a reasonable monetary estimate when only share quantities are recorded.

Final Validation and Error Handling

After applying placeholders or computed values, the predictor performs a strict NaN check across all six required columns. If any null values persist in open, high, low, close, volume, or amount, the code raises a ValueError to halt execution before malformed data reaches the inference engine.

Code Examples

DataFrame Missing Both Volume and Amount

import pandas as pd
from model.kronos import KronosPredictor, Kronos, KronosTokenizer

df = pd.DataFrame({
    "open":  [100, 101, 102],
    "high":  [101, 102, 103],
    "low":   [ 99, 100, 101],
    "close": [100.5, 101.5, 102.5]
    # No 'volume' or 'amount' columns

})

predictor = KronosPredictor(
    model=Kronos(...),               # model initialisation omitted for brevity

    tokenizer=KronosTokenizer(...)
)

# The predictor automatically adds volume=0 and amount=0

pred = predictor.predict(df, x_timestamp, y_timestamp, pred_len=5)
print(pred.head())

DataFrame with Volume but Missing Amount

df2 = pd.DataFrame({
    "open":  [200, 201, 202],
    "high":  [201, 202, 203],
    "low":   [199, 200, 201],
    "close": [200.5, 201.5, 202.5],
    "volume": [1500, 1600, 1550]      # volume provided

    # amount column omitted

})

# `amount` will be computed as volume * mean(price)

pred2 = predictor.predict(df2, x_timestamp, y_timestamp, pred_len=5)
print(pred2.head())

Source File Reference

The complete preprocessing logic resides in model/kronos.py within the KronosPredictor class, specifically inside the predict and predict_batch methods. Regression tests verifying these edge cases are located in tests/test_kronos_regression.py, while examples/prediction_example.py demonstrates typical usage patterns.

Summary

  • Automatic detection occurs inside predict() and predict_batch() methods, checking for self.vol_col and self.amt_vol.
  • Zero imputation fills both columns with 0.0 when they are completely absent from the input.
  • Smart derivation calculates amount = volume × mean(price) when volume exists but amount is missing.
  • Strict validation raises ValueError if NaN values remain after repair, protecting model integrity.

Frequently Asked Questions

What happens if my DataFrame has volume data but is missing the amount column?

Kronos computes the amount column automatically using the formula amount = volume × mean([open, high, low, close]). This provides a monetary turnover estimate based on the average price for each timestamp. If both columns are missing, they are both filled with zeros instead.

Does Kronos raise an error if volume and amount data are completely absent?

No. When Kronos detects missing volume and amount data, it creates both columns and fills them with 0.0 values. An error only occurs if NaN values persist after this automatic repair, or if the core OHLC price columns contain missing data.

Where is the missing data handling logic implemented in the Kronos source code?

The logic resides in model/kronos.py inside the KronosPredictor class. Specifically, the predict and predict_batch methods contain the column detection, placeholder insertion, amount derivation, and final validation steps. Tests covering these scenarios exist in tests/test_kronos_regression.py.

Can I manually provide amount values while omitting volume data?

The current implementation prioritizes volume presence when deriving amount. If volume is missing, both fields become zero placeholders. To use custom amount values without volume, you must provide both columns explicitly, as the automatic derivation only triggers when volume exists and amount is absent.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →