How Kronos TemporalEmbedding Encodes Time Information: Component-Based Temporal Encoding
Kronos TemporalEmbedding encodes time information by decomposing each timestamp into five discrete calendar components—minute, hour, weekday, day of month, and month—then embedding each component individually and summing the results into a unified temporal vector that is added to the model's token embeddings.
The shiyu-coder/Kronos repository implements a time-series transformer architecture that requires explicit temporal context for accurate sequence modeling. The Kronos TemporalEmbedding module provides this capability by converting raw timestamps into structured, multi-resolution representations that capture cyclical patterns at different granularities, from hourly cycles to monthly seasonality.
Decomposing Timestamps with calc_time_stamps
Before embedding occurs, raw timestamps must be converted into discrete categorical indices. The calc_time_stamps function in model/kronos.py handles this preprocessing by extracting five distinct integer fields from a pandas.Series of datetime objects.
# model/kronos.py
def calc_time_stamps(x_timestamp):
time_df = pd.DataFrame()
time_df['minute'] = x_timestamp.dt.minute
time_df['hour'] = x_timestamp.dt.hour
time_df['weekday'] = x_timestamp.dt.weekday
time_df['day'] = x_timestamp.dt.day
time_df['month'] = x_timestamp.dt.month
return time_df
This decomposition allows the model to learn distinct patterns for different temporal scales, such as intraday volatility (minute/hour), weekly seasonality (weekday), and monthly cycles (day/month).
Component-wise Embedding Strategy
The TemporalEmbedding class defined in model/module.py creates a separate embedding matrix for each of the five temporal components. Each embedding projects its respective time field into the model's hidden dimension (d_model).
Fixed Sinusoidal vs. Learnable Embeddings
According to the source code, the embedding type depends on the learn_pe parameter passed during initialization. When learn_pe=False, the module uses FixedEmbedding (sinusoidal position encodings), providing inductive biases for temporal cyclicality. When learn_pe=True, standard PyTorch nn.Embedding layers allow the model to learn temporal representations from data.
# model/module.py
class TemporalEmbedding(nn.Module):
def __init__(self, d_model, learn_pe):
super(TemporalEmbedding, self).__init__()
minute_size = 60
hour_size = 24
weekday_size = 7
day_size = 32 # 1-31 + padding
month_size = 13 # 1-12 + padding
Embed = FixedEmbedding if not learn_pe else nn.Embedding
self.minute_embed = Embed(minute_size, d_model)
self.hour_embed = Embed(hour_size, d_model)
self.weekday_embed = Embed(weekday_size, d_model)
self.day_embed = Embed(day_size, d_model)
self.month_embed = Embed(month_size, d_model)
def forward(self, x):
x = x.long() # [batch, seq_len, 5]
minute = self.minute_embed (x[:, :, 0])
hour = self.hour_embed (x[:, :, 1])
weekday = self.weekday_embed(x[:, :, 2])
day = self.day_embed (x[:, :, 3])
month = self.month_embed (x[:, :, 4])
# Summation yields a single temporal embedding per token
return hour + weekday + day + month + minute
The embedding dimensions correspond to the maximum values for each field: 60 possible minutes, 24 hours, 7 weekdays, 32 days (including padding), and 13 months (including padding).
Fusing Temporal and Token Embeddings
In the main Kronos class within model/kronos.py, the temporal embedding is instantiated and integrated into the forward pass. The module adds the temporal vector to the token embeddings, effectively providing positional and calendar context to each time-series token.
# model/kronos.py (excerpt)
self.time_emb = TemporalEmbedding(self.d_model, self.learn_te)
def forward(self, s1_ids, s2_ids, stamp=None, ...):
x = self.embedding([s1_ids, s2_ids]) # token embeddings
if stamp is not None:
time_embedding = self.time_emb(stamp) # temporal encoding
x = x + time_embedding
...
The stamp tensor supplied to forward must contain the five integer fields (minute, hour, weekday, day, month) for each token in the sequence, resulting in a shape of [batch_size, sequence_length, 5].
Practical Implementation Example
Here is a complete workflow demonstrating how to prepare temporal data and pass it through the model:
import pandas as pd
import torch
from model.kronos import Kronos, calc_time_stamps
from transformers import AutoTokenizer # assuming KronosTokenizer is HF-style
# 1️⃣ Load a pretrained Kronos model
model = Kronos.from_pretrained("NeoQuasar/Kronos-small")
tokenizer = ... # KronosTokenizer instance
# 2️⃣ Prepare data
price_df = pd.read_csv("data/example.csv", parse_dates=["timestamps"])
x_prices = price_df[["open","high","low","close","volume"]].values.astype("float32")
x_timestamp = price_df["timestamps"]
# 3️⃣ Convert timestamps to the 5-field integer matrix
time_df = calc_time_stamps(x_timestamp) # DataFrame with minute-hour-…
stamp_tensor = torch.tensor(time_df.values).unsqueeze(0) # shape [1, seq_len, 5]
# 4️⃣ Tokenise the price series (illustrative)
s1_ids, s2_ids = tokenizer.encode(price_df) # shape [1, seq_len]
# 5️⃣ Run the model with temporal information
s1_logits, s2_logits = model(s1_ids, s2_ids, stamp=stamp_tensor)
Summary
- Kronos TemporalEmbedding processes time by splitting timestamps into five discrete components: minute, hour, weekday, day, and month.
- Each component is embedded independently using either fixed sinusoidal patterns or learnable embeddings, controlled by the
learn_peparameter. - The component embeddings are summed to produce a single temporal vector per sequence position.
- This temporal vector is added to the token embeddings in the main
Kronosforward pass, as implemented inmodel/kronos.py. - The input
stamptensor must have shape[batch, seq_len, 5]and contain integer indices for each temporal field.
Frequently Asked Questions
What are the five temporal components used in Kronos TemporalEmbedding?
The five components are minute (0-59), hour (0-23), weekday (0-6), day of month (1-31), and month (1-12). These are extracted by the calc_time_stamps function in model/kronos.py and correspond to indices 0 through 4 in the embedding lookup.
How does the learn_pe parameter affect temporal encoding?
When learn_pe=False, the model uses FixedEmbedding (sinusoidal encodings) which provide built-in inductive biases for cyclical time patterns. When learn_pe=True, the model uses standard nn.Embedding layers that learn temporal representations from scratch during training, allowing for data-specific temporal patterns.
Why does Kronos sum the component embeddings instead of concatenating them?
Summing the embeddings (minute + hour + weekday + day + month) ensures the output remains in the same dimensional space as the token embeddings (d_model), allowing simple addition for fusion. Concatenation would increase dimensionality and require additional projection layers, whereas summation maintains efficiency and allows the model to learn composite temporal representations directly.
What input shape does the stamp parameter require in the Kronos forward method?
The stamp parameter requires a tensor of shape [batch_size, sequence_length, 5] containing integer indices. The last dimension must follow the specific order: minute (index 0), hour (1), weekday (2), day (3), and month (4), as defined by the calc_time_stamps function output.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →