# How Temperature, Top‑k, and Top‑p Parameters Affect Sampling in Kronos Probabilistic Forecasting

> Discover how temperature, top-k, and top-p parameters influence sampling in Kronos probabilistic forecasting. Control distribution sharpness and limit vocabulary for stochastic price-volume forecasts.

- Repository: [ShiYu/Kronos](https://github.com/shiyu-coder/Kronos)
- Tags: deep-dive
- Published: 2026-04-10

---

**In Kronos, temperature scales the logits to control distribution sharpness, while top‑k and top‑p limit the sampling vocabulary to the most probable tokens, jointly determining the stochasticity of autoregressive price‑volume forecasts.**

Kronos is an open‑source Transformer‑based model for financial time‑series forecasting that generates future sequences through **autoregressive sampling**. The **temperature, top_k, and top_p parameters** control the randomness and diversity of these predictions by modifying the probability distribution over discrete tokens at each decoding step. Understanding how these three mechanisms interact is essential for calibrating the trade‑off between deterministic accuracy and exploratory scenario generation in probabilistic forecasting.

## Temperature Scaling

### Implementation in sample_from_logits

In [`model/kronos.py`](https://github.com/shiyu-coder/Kronos/blob/main/model/kronos.py), the `sample_from_logits` function (lines 73‑77) applies **temperature scaling** by dividing the raw logits by a scalar `temperature` value before computing the softmax:

```python
logits = logits / temperature

```

This linear transformation occurs immediately before the softmax operation, directly influencing the entropy of the resulting categorical distribution. The scaled logits are then passed to `top_k_top_p_filtering` before the final multinomial sampling draw.

### Impact on Forecast Diversity

**High temperature (> 1)** flattens the probability distribution, increasing the likelihood of sampling lower‑probability tokens and producing more diverse but potentially noisier forecasts. **Low temperature (< 1)** sharpens the distribution, concentrating probability mass on the highest‑scoring tokens and yielding more deterministic, conservative predictions. At **temperature = 1.0**, the original logits remain unmodified, serving as the neutral baseline for standard sampling behavior.

## Top‑k Filtering

### Fixed Vocabulary Truncation

The `top_k_top_p_filtering` function in [`model/kronos.py`](https://github.com/shiyu-coder/Kronos/blob/main/model/kronos.py) (lines 31‑46) implements **top‑k filtering** by retaining only the *k* highest‑scoring tokens and replacing the logits of all other tokens with negative infinity (`‑inf`). This hard truncation eliminates extreme outliers from the candidate pool, ensuring the model selects only from a fixed‑size shortlist of the most probable market movements at each time step.

### Configuration and Disablement

When `top_k` is set to `0`, the filter is disabled entirely, allowing the full vocabulary to participate in sampling. Typical production values range from 10 to 100, creating a constrained candidate set that limits variance while preserving the most likely forecast trajectories.

## Top‑p (Nucleus) Sampling

### Dynamic Threshold Mechanism

Also implemented within `top_k_top_p_filtering` (lines 54‑70), **top‑p sampling**—or nucleus sampling—retains the smallest possible set of tokens whose cumulative probability exceeds the threshold *p*. The algorithm sorts the temperature‑scaled logits, computes cumulative probabilities using softmax, and masks any token beyond the cutoff with the filter value, ensuring only high‑cumulative‑probability tokens remain eligible.

### Adaptive Vocabulary Reduction

Unlike the fixed count of top‑k, top‑p adapts dynamically to the distribution’s shape: peaked distributions yield small candidate sets, while flat distributions retain more tokens. This provides a **dynamic shortlist** that maintains diversity when model uncertainty is high but enforces strict quality control when confidence is concentrated in a few dominant tokens.

## Practical Configuration and Code Examples

Parameters flow from the high‑level `KronosPredictor.predict` method through `auto_regressive_inference` (lines 438‑445 in [`model/kronos.py`](https://github.com/shiyu-coder/Kronos/blob/main/model/kronos.py)) to the sampler, where they are forwarded to `sample_from_logits`. The following examples demonstrate how to configure **temperature**, **top_k**, and **top_p** for different forecasting objectives:

```python
import pandas as pd
from kronos import KronosPredictor, Kronos, KronosTokenizer

# Assume `model` and `tokenizer` are already loaded (e.g. via torch.load)

predictor = KronosPredictor(model, tokenizer)

# Historical price/volume data (DataFrame with open, high, low, close columns)

df = pd.read_csv("examples/data/XSHG_5min_600977.csv", parse_dates=["timestamp"])
x_ts = df["timestamp"]
y_ts = pd.date_range(start=x_ts.iloc[-1] + pd.Timedelta(minutes=5),
                     periods=20, freq="5min")

# ------------------------------

# 1️⃣ Deterministic forecast (low temperature, no sampling filters)

# ------------------------------

deterministic = predictor.predict(
    df,
    x_timestamp=x_ts,
    y_timestamp=y_ts,
    pred_len=20,
    T=0.5,          # low temperature → sharper distribution

    top_k=0,        # disabled

    top_p=0.99,     # practically disabled

    sample_count=1,
    verbose=False,
)

# ------------------------------

# 2️⃣ Diverse stochastic forecast (high temperature, nucleus sampling)

# ------------------------------

stochastic = predictor.predict(
    df,
    x_timestamp=x_ts,
    y_timestamp=y_ts,
    pred_len=20,
    T=2.0,          # high temperature → flatter distribution

    top_k=0,        # keep nucleus sampling only

    top_p=0.8,      # keep only tokens covering 80 % probability mass

    sample_count=5, # average over 5 Monte‑Carlo samples

    verbose=False,
)

# ------------------------------

# 3️⃣ Fixed‑size shortlist (top‑k filtering)

# ------------------------------

topk_forecast = predictor.predict(
    df,
    x_timestamp=x_ts,
    y_timestamp=y_ts,
    pred_len=20,
    T=1.0,
    top_k=50,       # keep only the 50 most likely tokens at each step

    top_p=1.0,      # disable top‑p

    sample_count=1,
    verbose=False,
)

print("Deterministic head:", deterministic.head())
print("Stochastic head:", stochastic.head())
print("Top‑k head:", topk_forecast.head())

```

These configurations illustrate how **low temperature + strong constraints** produce risk‑averse, deterministic forecasts, while **high temperature + weak or no constraints** generate exploratory scenarios suitable for stress testing and Monte Carlo simulation.

## Summary

- **Temperature** scales logits directly in `sample_from_logits` to control distribution entropy: values < 1 sharpen predictions, values > 1 increase diversity.
- **Top‑k** enforces a fixed‑size vocabulary of the *k* most likely tokens in `top_k_top_p_filtering`, explicitly capping the candidate pool and preventing outlier selection.
- **Top‑p** uses a dynamic cumulative probability threshold in the same filtering function, adapting the vocabulary size to the model’s confidence at each autoregressive step.
- All three parameters are optional and composable; they are forwarded from `KronosPredictor` through `auto_regressive_inference` (lines 438‑445) to control the stochasticity of probabilistic forecasts.

## Frequently Asked Questions

### What happens if I set both top_k and top_p simultaneously in Kronos?

When both parameters are provided, the `top_k_top_p_filtering` function applies both constraints sequentially: it first reduces the logits to the top‑k candidates, then further filters those results using the top‑p cumulative threshold. This intersection ensures the sampling pool satisfies both the fixed count and the cumulative probability requirements before multinomial sampling occurs.

### How does temperature differ from top‑k and top‑p in Kronos forecasting?

**Temperature** modifies the probability distribution’s shape globally before any filtering occurs, affecting the relative likelihood of all tokens continuously. In contrast, **top‑k** and **top‑p** perform hard truncation on the vocabulary set, completely eliminating certain tokens from consideration by setting their logits to `‑inf`, regardless of the temperature scaling applied earlier.

### Why would I use high temperature with strong top‑k constraints?

This combination allows you to explore diverse trajectories (via high temperature flattening the distribution) while preventing extreme outliers (via top‑k truncation). The model considers unlikely but plausible alternatives to the most probable path, yet avoids sampling tokens that rank below the top‑k cutoff, effectively balancing exploration with risk management in volatile market forecasts.

### Where are the default sampling values defined in the Kronos repository?

Default values are typically set in the `KronosPredictor` class methods and the `auto_regressive_inference` function within [`model/kronos.py`](https://github.com/shiyu-coder/Kronos/blob/main/model/kronos.py). The common defaults are `temperature=1.0`, `top_k=0` (disabled), and `top_p=0.99`, though these can be overridden via the `predict` and `predict_batch` method arguments. Additionally, the Web UI in [`webui/app.py`](https://github.com/shiyu-coder/Kronos/blob/main/webui/app.py) exposes these parameters for interactive adjustment.