Batch Prediction Consistency Requirements in Kronos: A Complete Guide to predict_batch()

To ensure batch prediction consistency in Kronos, all input series must share identical historical lengths, identical prediction horizons, and contain the required OHLCV columns without NaN values, while the three input lists themselves must be equal-length iterables.

The predict_batch() method in the shiyu-coder/Kronos repository enables high-throughput parallel inference across multiple time-series simultaneously. Maintaining strict batch prediction consistency is essential for the method to correctly stack tensors and execute autoregressive generation on GPU. The implementation enforces these constraints through a rigorous validation pipeline defined in model/kronos.py.

Input Container Validation

Before processing any market data, predict_batch() verifies the structure of its three primary arguments. According to lines 562-586 in model/kronos.py, the method requires that df_list, x_timestamp_list, and y_timestamp_list are either lists or tuples, and that they share identical lengths:

if not isinstance(df_list, (list, tuple)) ...
if not (len(df_list) == len(x_timestamp_list) == len(y_timestamp_list)):
    raise ValueError(...)

This check guarantees a one-to-one mapping between historical data, historical timestamps, and future prediction timestamps across the entire batch.

Per-Series Schema Requirements

For each element in df_list, the method performs strict type and content validation (lines 597-613 in model/kronos.py). Every item must be a pandas.DataFrame containing the five mandatory price-related columns: open, high, low, close, volume, and amount.

The validation pipeline automatically handles missing volume or amount columns by filling them with zeros or derived values. However, any NaN values present in the price or volume fields after this preprocessing will trigger an immediate error, preventing undefined numeric operations during normalization.

Temporal Alignment and Length Constraints

Strict temporal consistency rules govern the historical and prediction horizons. After normalizing each series, the method records the historical length (seq_lens) and prediction length (y_lens) for every item in the batch. As implemented in lines 642-647 of model/kronos.py, all series must share exactly the same historical length and exactly the same prediction length (which must equal the pred_len argument):


# Failure to meet either condition raises:

# ValueError: Parallel prediction requires all series to have consistent historical lengths

Additionally, the method converts raw timestamps into five time-feature columns (minute, hour, weekday, day, month) using the calc_time_stamps helper function (lines 72-79). These temporal features must align perfectly with the corresponding historical or prediction windows.

Tensor Construction and Parallel Generation

Once length consistency is confirmed, the normalized price/volume arrays and time-feature arrays are stacked along a new batch dimension. This produces tensors with the following shapes:

  • x_batch → (B, seq_len, feat)
  • x_stamp_batch → (B, seq_len, time_feat)
  • y_stamp_batch → (B, pred_len, time_feat)

These stacked tensors are passed to the internal generate() method (lines 652-655), which executes autoregressive inference once for the entire batch. After generation, predictions are de-scaled using per-series mean and standard deviation values stored during preprocessing, then wrapped in pandas.DataFrame objects indexed by the supplied y_timestamp values.

Practical Implementation

The following example demonstrates valid batch prediction where all consistency requirements are satisfied:

import pandas as pd
from model import Kronos, KronosTokenizer, KronosPredictor

# Load model and tokenizer

tokenizer = KronosTokenizer.from_pretrained('.../Kronos-Tokenizer-base/')
model = Kronos.from_pretrained('.../Kronos-base/')
predictor = KronosPredictor(model, tokenizer, device='cuda:0', max_context=512)

# Prepare a list of 5 historical windows (all length 400) and matching timestamps

dfs, x_ts, y_ts = [], [], []
lookback, pred_len = 400, 120
for i in range(5):
    hist = df.loc[i*400:(i*400+lookback-1),
                 ['open','high','low','close','volume','amount']]
    dfs.append(hist)
    x_ts.append(df.loc[i*400:(i*400+lookback-1), 'timestamps'])
    y_ts.append(df.loc[i*400+lookback:i*400+lookback+pred_len-1, 'timestamps'])

# Batch prediction – all series share the same 400‑step history and 120‑step horizon

pred_dfs = predictor.predict_batch(
    df_list=dfs,
    x_timestamp_list=x_ts,
    y_timestamp_list=y_ts,
    pred_len=pred_len,
)

# `pred_dfs` is a list of DataFrames, each indexed by its future timestamps

Attempting to batch series with mismatched lengths triggers the explicit validation error defined in the source:


# Mismatched historical lengths → raises ValueError

short_hist = df.iloc[:300][['open','high','low','close','volume','amount']]
predictor.predict_batch(
    df_list=[hist, short_hist],          # <-- different seq_len

    x_timestamp_list=[x_ts[0], x_ts[1][:300]],
    y_timestamp_list=[y_ts[0], y_ts[1]],
    pred_len=120,
)

# -> ValueError: Parallel prediction requires all series to have consistent historical lengths

Summary

  • Input containers must be equal-length lists or tuples to ensure one-to-one mapping of data and timestamps.
  • Per-series data must be valid pandas.DataFrame objects containing OHLCV columns with no NaN values.
  • Length uniformity is mandatory: every series in the batch must share the identical seq_len and identical pred_len to enable tensor stacking.
  • Temporal features are derived via calc_time_stamps and must align with the historical and prediction windows.
  • GPU efficiency relies on the batch tensor format (B, seq_len, feat), which requires strict homogeneity across the batch dimension.

Frequently Asked Questions

What happens if my time-series have different historical lengths?

The method raises a ValueError immediately. According to lines 642-647 in model/kronos.py, parallel prediction requires all series to have consistent historical lengths to allow stacking into a single tensor for GPU inference. You must either truncate longer series or pad shorter ones to a common seq_len before calling predict_batch().

Which columns are mandatory in the input DataFrames?

Each DataFrame must contain the open, high, low, close, volume, and amount columns. The validation logic in lines 597-613 of model/kronos.py checks for these specifically. Missing volume or amount data is automatically filled with zeros, but missing price columns or NaN values in any critical field will cause the method to fail.

How does predict_batch() handle timestamps?

The method converts raw timestamps into five categorical features (minute, hour, weekday, day, month) using the internal calc_time_stamps function (lines 72-79). These features are concatenated with price data to form the final input tensors. The x_timestamp_list provides features for the historical window, while y_timestamp_list provides features for the autoregressive generation horizon.

Can I mix different prediction lengths in a single batch call?

No. All series must share the same pred_len parameter. The autoregressive generate() method runs once for the entire batch and requires a consistent future horizon to manage the decoder's temporal attention and output dimensions. Attempting to mix lengths will trigger the same validation error as mismatched historical lengths.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →