How Temperature, Top‑k, and Top‑p Parameters Affect Sampling in Kronos Probabilistic Forecasting
In Kronos, temperature scales the logits to control distribution sharpness, while top‑k and top‑p limit the sampling vocabulary to the most probable tokens, jointly determining the stochasticity of autoregressive price‑volume forecasts.
Kronos is an open‑source Transformer‑based model for financial time‑series forecasting that generates future sequences through autoregressive sampling. The temperature, top_k, and top_p parameters control the randomness and diversity of these predictions by modifying the probability distribution over discrete tokens at each decoding step. Understanding how these three mechanisms interact is essential for calibrating the trade‑off between deterministic accuracy and exploratory scenario generation in probabilistic forecasting.
Temperature Scaling
Implementation in sample_from_logits
In model/kronos.py, the sample_from_logits function (lines 73‑77) applies temperature scaling by dividing the raw logits by a scalar temperature value before computing the softmax:
logits = logits / temperature
This linear transformation occurs immediately before the softmax operation, directly influencing the entropy of the resulting categorical distribution. The scaled logits are then passed to top_k_top_p_filtering before the final multinomial sampling draw.
Impact on Forecast Diversity
High temperature (> 1) flattens the probability distribution, increasing the likelihood of sampling lower‑probability tokens and producing more diverse but potentially noisier forecasts. Low temperature (< 1) sharpens the distribution, concentrating probability mass on the highest‑scoring tokens and yielding more deterministic, conservative predictions. At temperature = 1.0, the original logits remain unmodified, serving as the neutral baseline for standard sampling behavior.
Top‑k Filtering
Fixed Vocabulary Truncation
The top_k_top_p_filtering function in model/kronos.py (lines 31‑46) implements top‑k filtering by retaining only the k highest‑scoring tokens and replacing the logits of all other tokens with negative infinity (‑inf). This hard truncation eliminates extreme outliers from the candidate pool, ensuring the model selects only from a fixed‑size shortlist of the most probable market movements at each time step.
Configuration and Disablement
When top_k is set to 0, the filter is disabled entirely, allowing the full vocabulary to participate in sampling. Typical production values range from 10 to 100, creating a constrained candidate set that limits variance while preserving the most likely forecast trajectories.
Top‑p (Nucleus) Sampling
Dynamic Threshold Mechanism
Also implemented within top_k_top_p_filtering (lines 54‑70), top‑p sampling—or nucleus sampling—retains the smallest possible set of tokens whose cumulative probability exceeds the threshold p. The algorithm sorts the temperature‑scaled logits, computes cumulative probabilities using softmax, and masks any token beyond the cutoff with the filter value, ensuring only high‑cumulative‑probability tokens remain eligible.
Adaptive Vocabulary Reduction
Unlike the fixed count of top‑k, top‑p adapts dynamically to the distribution’s shape: peaked distributions yield small candidate sets, while flat distributions retain more tokens. This provides a dynamic shortlist that maintains diversity when model uncertainty is high but enforces strict quality control when confidence is concentrated in a few dominant tokens.
Practical Configuration and Code Examples
Parameters flow from the high‑level KronosPredictor.predict method through auto_regressive_inference (lines 438‑445 in model/kronos.py) to the sampler, where they are forwarded to sample_from_logits. The following examples demonstrate how to configure temperature, top_k, and top_p for different forecasting objectives:
import pandas as pd
from kronos import KronosPredictor, Kronos, KronosTokenizer
# Assume `model` and `tokenizer` are already loaded (e.g. via torch.load)
predictor = KronosPredictor(model, tokenizer)
# Historical price/volume data (DataFrame with open, high, low, close columns)
df = pd.read_csv("examples/data/XSHG_5min_600977.csv", parse_dates=["timestamp"])
x_ts = df["timestamp"]
y_ts = pd.date_range(start=x_ts.iloc[-1] + pd.Timedelta(minutes=5),
periods=20, freq="5min")
# ------------------------------
# 1️⃣ Deterministic forecast (low temperature, no sampling filters)
# ------------------------------
deterministic = predictor.predict(
df,
x_timestamp=x_ts,
y_timestamp=y_ts,
pred_len=20,
T=0.5, # low temperature → sharper distribution
top_k=0, # disabled
top_p=0.99, # practically disabled
sample_count=1,
verbose=False,
)
# ------------------------------
# 2️⃣ Diverse stochastic forecast (high temperature, nucleus sampling)
# ------------------------------
stochastic = predictor.predict(
df,
x_timestamp=x_ts,
y_timestamp=y_ts,
pred_len=20,
T=2.0, # high temperature → flatter distribution
top_k=0, # keep nucleus sampling only
top_p=0.8, # keep only tokens covering 80 % probability mass
sample_count=5, # average over 5 Monte‑Carlo samples
verbose=False,
)
# ------------------------------
# 3️⃣ Fixed‑size shortlist (top‑k filtering)
# ------------------------------
topk_forecast = predictor.predict(
df,
x_timestamp=x_ts,
y_timestamp=y_ts,
pred_len=20,
T=1.0,
top_k=50, # keep only the 50 most likely tokens at each step
top_p=1.0, # disable top‑p
sample_count=1,
verbose=False,
)
print("Deterministic head:", deterministic.head())
print("Stochastic head:", stochastic.head())
print("Top‑k head:", topk_forecast.head())
These configurations illustrate how low temperature + strong constraints produce risk‑averse, deterministic forecasts, while high temperature + weak or no constraints generate exploratory scenarios suitable for stress testing and Monte Carlo simulation.
Summary
- Temperature scales logits directly in
sample_from_logitsto control distribution entropy: values < 1 sharpen predictions, values > 1 increase diversity. - Top‑k enforces a fixed‑size vocabulary of the k most likely tokens in
top_k_top_p_filtering, explicitly capping the candidate pool and preventing outlier selection. - Top‑p uses a dynamic cumulative probability threshold in the same filtering function, adapting the vocabulary size to the model’s confidence at each autoregressive step.
- All three parameters are optional and composable; they are forwarded from
KronosPredictorthroughauto_regressive_inference(lines 438‑445) to control the stochasticity of probabilistic forecasts.
Frequently Asked Questions
What happens if I set both top_k and top_p simultaneously in Kronos?
When both parameters are provided, the top_k_top_p_filtering function applies both constraints sequentially: it first reduces the logits to the top‑k candidates, then further filters those results using the top‑p cumulative threshold. This intersection ensures the sampling pool satisfies both the fixed count and the cumulative probability requirements before multinomial sampling occurs.
How does temperature differ from top‑k and top‑p in Kronos forecasting?
Temperature modifies the probability distribution’s shape globally before any filtering occurs, affecting the relative likelihood of all tokens continuously. In contrast, top‑k and top‑p perform hard truncation on the vocabulary set, completely eliminating certain tokens from consideration by setting their logits to ‑inf, regardless of the temperature scaling applied earlier.
Why would I use high temperature with strong top‑k constraints?
This combination allows you to explore diverse trajectories (via high temperature flattening the distribution) while preventing extreme outliers (via top‑k truncation). The model considers unlikely but plausible alternatives to the most probable path, yet avoids sampling tokens that rank below the top‑k cutoff, effectively balancing exploration with risk management in volatile market forecasts.
Where are the default sampling values defined in the Kronos repository?
Default values are typically set in the KronosPredictor class methods and the auto_regressive_inference function within model/kronos.py. The common defaults are temperature=1.0, top_k=0 (disabled), and top_p=0.99, though these can be overridden via the predict and predict_batch method arguments. Additionally, the Web UI in webui/app.py exposes these parameters for interactive adjustment.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →