# Batch Prediction Consistency Requirements in Kronos: A Complete Guide to predict_batch()

> Master Kronos batch prediction consistency with predict_batch(). Ensure identical series lengths, horizons, and OHLCV columns for accurate results. Learn the complete requirements.

- Repository: [ShiYu/Kronos](https://github.com/shiyu-coder/Kronos)
- Tags: deep-dive
- Published: 2026-04-10

---

**To ensure batch prediction consistency in Kronos, all input series must share identical historical lengths, identical prediction horizons, and contain the required OHLCV columns without NaN values, while the three input lists themselves must be equal-length iterables.**

The `predict_batch()` method in the [shiyu-coder/Kronos](https://github.com/shiyu-coder/Kronos) repository enables high-throughput parallel inference across multiple time-series simultaneously. Maintaining strict **batch prediction consistency** is essential for the method to correctly stack tensors and execute autoregressive generation on GPU. The implementation enforces these constraints through a rigorous validation pipeline defined in [`model/kronos.py`](https://github.com/shiyu-coder/Kronos/blob/main/model/kronos.py).

## Input Container Validation

Before processing any market data, `predict_batch()` verifies the structure of its three primary arguments. According to lines 562-586 in [`model/kronos.py`](https://github.com/shiyu-coder/Kronos/blob/main/model/kronos.py), the method requires that `df_list`, `x_timestamp_list`, and `y_timestamp_list` are either lists or tuples, and that they share identical lengths:

```python
if not isinstance(df_list, (list, tuple)) ...
if not (len(df_list) == len(x_timestamp_list) == len(y_timestamp_list)):
    raise ValueError(...)

```

This check guarantees a one-to-one mapping between historical data, historical timestamps, and future prediction timestamps across the entire batch.

## Per-Series Schema Requirements

For each element in `df_list`, the method performs strict type and content validation (lines 597-613 in [`model/kronos.py`](https://github.com/shiyu-coder/Kronos/blob/main/model/kronos.py)). Every item must be a `pandas.DataFrame` containing the five mandatory price-related columns: **open**, **high**, **low**, **close**, **volume**, and **amount**.

The validation pipeline automatically handles missing volume or amount columns by filling them with zeros or derived values. However, any `NaN` values present in the price or volume fields after this preprocessing will trigger an immediate error, preventing undefined numeric operations during normalization.

## Temporal Alignment and Length Constraints

Strict **temporal consistency** rules govern the historical and prediction horizons. After normalizing each series, the method records the historical length (`seq_lens`) and prediction length (`y_lens`) for every item in the batch. As implemented in lines 642-647 of [`model/kronos.py`](https://github.com/shiyu-coder/Kronos/blob/main/model/kronos.py), all series must share exactly the same historical length and exactly the same prediction length (which must equal the `pred_len` argument):

```python

# Failure to meet either condition raises:

# ValueError: Parallel prediction requires all series to have consistent historical lengths

```

Additionally, the method converts raw timestamps into five time-feature columns (**minute**, **hour**, **weekday**, **day**, **month**) using the `calc_time_stamps` helper function (lines 72-79). These temporal features must align perfectly with the corresponding historical or prediction windows.

## Tensor Construction and Parallel Generation

Once length consistency is confirmed, the normalized price/volume arrays and time-feature arrays are stacked along a new batch dimension. This produces tensors with the following shapes:

- `x_batch` → **(B, seq_len, feat)**
- `x_stamp_batch` → **(B, seq_len, time_feat)**
- `y_stamp_batch` → **(B, pred_len, time_feat)**

These stacked tensors are passed to the internal `generate()` method (lines 652-655), which executes autoregressive inference once for the entire batch. After generation, predictions are de-scaled using per-series mean and standard deviation values stored during preprocessing, then wrapped in `pandas.DataFrame` objects indexed by the supplied `y_timestamp` values.

## Practical Implementation

The following example demonstrates valid batch prediction where all consistency requirements are satisfied:

```python
import pandas as pd
from model import Kronos, KronosTokenizer, KronosPredictor

# Load model and tokenizer

tokenizer = KronosTokenizer.from_pretrained('.../Kronos-Tokenizer-base/')
model = Kronos.from_pretrained('.../Kronos-base/')
predictor = KronosPredictor(model, tokenizer, device='cuda:0', max_context=512)

# Prepare a list of 5 historical windows (all length 400) and matching timestamps

dfs, x_ts, y_ts = [], [], []
lookback, pred_len = 400, 120
for i in range(5):
    hist = df.loc[i*400:(i*400+lookback-1),
                 ['open','high','low','close','volume','amount']]
    dfs.append(hist)
    x_ts.append(df.loc[i*400:(i*400+lookback-1), 'timestamps'])
    y_ts.append(df.loc[i*400+lookback:i*400+lookback+pred_len-1, 'timestamps'])

# Batch prediction – all series share the same 400‑step history and 120‑step horizon

pred_dfs = predictor.predict_batch(
    df_list=dfs,
    x_timestamp_list=x_ts,
    y_timestamp_list=y_ts,
    pred_len=pred_len,
)

# `pred_dfs` is a list of DataFrames, each indexed by its future timestamps

```

Attempting to batch series with mismatched lengths triggers the explicit validation error defined in the source:

```python

# Mismatched historical lengths → raises ValueError

short_hist = df.iloc[:300][['open','high','low','close','volume','amount']]
predictor.predict_batch(
    df_list=[hist, short_hist],          # <-- different seq_len

    x_timestamp_list=[x_ts[0], x_ts[1][:300]],
    y_timestamp_list=[y_ts[0], y_ts[1]],
    pred_len=120,
)

# -> ValueError: Parallel prediction requires all series to have consistent historical lengths

```

## Summary

- **Input containers** must be equal-length lists or tuples to ensure one-to-one mapping of data and timestamps.
- **Per-series data** must be valid `pandas.DataFrame` objects containing OHLCV columns with no `NaN` values.
- **Length uniformity** is mandatory: every series in the batch must share the identical `seq_len` and identical `pred_len` to enable tensor stacking.
- **Temporal features** are derived via `calc_time_stamps` and must align with the historical and prediction windows.
- **GPU efficiency** relies on the batch tensor format `(B, seq_len, feat)`, which requires strict homogeneity across the batch dimension.

## Frequently Asked Questions

### What happens if my time-series have different historical lengths?

The method raises a `ValueError` immediately. According to lines 642-647 in [`model/kronos.py`](https://github.com/shiyu-coder/Kronos/blob/main/model/kronos.py), parallel prediction requires all series to have consistent historical lengths to allow stacking into a single tensor for GPU inference. You must either truncate longer series or pad shorter ones to a common `seq_len` before calling `predict_batch()`.

### Which columns are mandatory in the input DataFrames?

Each DataFrame must contain the **open**, **high**, **low**, **close**, **volume**, and **amount** columns. The validation logic in lines 597-613 of [`model/kronos.py`](https://github.com/shiyu-coder/Kronos/blob/main/model/kronos.py) checks for these specifically. Missing volume or amount data is automatically filled with zeros, but missing price columns or NaN values in any critical field will cause the method to fail.

### How does `predict_batch()` handle timestamps?

The method converts raw timestamps into five categorical features (minute, hour, weekday, day, month) using the internal `calc_time_stamps` function (lines 72-79). These features are concatenated with price data to form the final input tensors. The `x_timestamp_list` provides features for the historical window, while `y_timestamp_list` provides features for the autoregressive generation horizon.

### Can I mix different prediction lengths in a single batch call?

No. All series must share the same `pred_len` parameter. The autoregressive `generate()` method runs once for the entire batch and requires a consistent future horizon to manage the decoder's temporal attention and output dimensions. Attempting to mix lengths will trigger the same validation error as mismatched historical lengths.