How Kronos Tokenizes Financial Candlestick Data into s1 and s2 Tokens
Kronos converts raw OHLCV candlestick data into a two-level hierarchical token stream using a Binary Spherical Quantizer that splits binary representations into high-order s1 (pre‑token) and low-order s2 (post‑token) indices.
Kronos is an open‑source financial time‑series model that represents market data as discrete tokens rather than continuous values. The tokenization process transforms normalized price‑volume tensors into compact binary codes that are split into s1 and s2 tokens, enabling the model to capture market structure at both coarse and fine granularities.
The Three-Stage Tokenization Pipeline
The tokenization flow implemented in shiyu‑coder/Kronos follows a clear pipeline from raw dataframe to paired token indices. According to the source code in model/kronos.py and model/module.py, the process involves normalization, binary‑spherical quantization, and hierarchical splitting.
Stage 1: Normalization and Tensor Preparation
Before quantization, the KronosPredictor normalizes the raw OHLCV (open‑high‑low‑close‑volume) dataframe and constructs a 3‑D tensor x with shape [batch, seq_len, feat]. This normalized tensor serves as the continuous input to the tokenizer, ensuring that price and volume scales do not bias the subsequent binary encoding.
Stage 2: Binary-Spherical Quantization
The core quantization logic resides in BSQuantizer (Binary Spherical Quantizer) defined in model/module.py. When KronosTokenizer.encode(..., half=True) is invoked, the normalized tensor flows through the following steps:
- L2 Normalization – Features are projected onto a unit sphere to remove magnitude variations.
- Binary Code Generation – The quantizer produces a binary code of length
codebook_dim(equal tos1 + s2bits). - Bit‑to‑Index Conversion – The
bits_to_indicesmethod interprets the binary vector as an integer using power‑of‑two weighting.
# model/module.py – BSQuantizer.bits_to_indices (lines 45-55 context)
def bits_to_indices(self, bits):
bits = (bits >= 0).to(torch.long)
indices = 2 ** torch.arange(
0,
bits.shape[-1],
1,
dtype=torch.long,
device=bits.device,
)
return (bits * indices).sum(-1)
When the half=True flag is passed to encode(), the full binary code is split into two contiguous segments: the first s1 bits and the last s2 bits. Each segment is converted to a token index tensor independently, yielding two separate outputs rather than a single composite ID.
Stage 3: Hierarchical Token Splitting
For scenarios where tokens are represented as a single composite integer (e.g., when loading pre‑computed token IDs), the HierarchicalEmbedding class provides a split_token method to decompose the integer back into s1 and s2 components using bitwise operations:
# model/module.py – HierarchicalEmbedding.split_token (lines 17-28)
def split_token(self, token_ids: torch.Tensor, s2_bits: int):
t = token_ids.long()
mask = (1 << s2_bits) - 1
s2_ids = t & mask # low bits → s₂
s1_ids = t >> s2_bits # high bits → s₁
return s1_ids, s2_ids
This utility allows the model to accept either pre‑split token pairs or unified tokens that get separated inside the embedding layer.
Practical Implementation: Tokenizing OHLCV Data
The following example demonstrates the end‑to‑end tokenization of a candlestick CSV using the public KronosTokenizer:
import pandas as pd
import torch
from model.kronos import KronosTokenizer
# Load OHLCV data
df = pd.read_csv("examples/data/XSHG_5min_600977.csv")
price_cols = ["open", "high", "low", "close"]
df = df[price_cols + ["volume"]]
# Initialize tokenizer (s₁=8 bits, s₂=8 bits in base model)
tokenizer = KronosTokenizer.from_pretrained(
"NeoQuasar/Kronos-Tokenizer-base", revision="v1.0"
)
# Prepare tensor: [batch=1, time, features]
x_tensor = torch.from_numpy(df.values.astype("float32")).unsqueeze(0)
# Encode with half=True to receive separate s1 and s2 tensors
s1_ids, s2_ids = tokenizer.encode(x_tensor, half=True)
print(f"s₁ shape: {s1_ids.shape}") # torch.Size([1, T])
print(f"s₂ shape: {s2_ids.shape}") # torch.Size([1, T])
The encode method in model/kronos.py (lines 42‑53) orchestrates this call, forwarding the tensor to BSQuantizer.forward and returning the split indices when half=True is specified.
Understanding s1 and s2 Token Roles
The hierarchical design separates market information into interdependent abstraction levels:
- s1 Tokens (Pre‑tokens) – Represented by the high‑order bits of the binary code, these act as a coarse vocabulary. The model first processes s1 tokens to establish a broad market context or "theme" for the current time step.
- s2 Tokens (Post‑tokens) – Represented by the low‑order bits, these provide fine‑grained detail. Within the transformer architecture, s2 generation is conditioned on the decoded s1 context through a dependency‑aware layer, ensuring that granular price movements remain grounded in the broader market structure established by s1.
This bit‑level separation allows Kronos to model the conditional distribution P(s2 | s1), effectively learning hierarchical dependencies in financial time series while maintaining a compact codebook of size 2^(s1+s2).
Summary
- Kronos tokenizes financial candlestick data via Binary Spherical Quantization in
model/module.py. - The
BSQuantizergenerates binary codes of lengths1 + s2bits via L2‑normalized projection. - Setting
half=TrueinKronosTokenizer.encode()splits the binary string into separate s1 (high‑order) and s2 (low‑order) token tensors. HierarchicalEmbedding.split_tokencan decompose composite token IDs using bitwise masks and shifts.- s1 captures coarse market structure while s2 refines it, enabling conditional generation within the model architecture.
Frequently Asked Questions
What do s1 and s2 stand for in Kronos tokenization?
s1 refers to "pre‑tokens" (high‑order bits) and s2 refers to "post‑tokens" (low‑order bits). This naming reflects the hierarchical generation order: the model first predicts or attends to s1 tokens to establish context, then generates s2 tokens conditioned on that context, similar to a coarse‑to‑fine quantization strategy.
Why does Kronos use binary quantization instead of learned vector quantization?
Binary Spherical Quantization maps normalized price movements to vertices of a hypercube inscribed on a unit sphere. This approach provides deterministic, reversible encoding without requiring a learned codebook, ensuring that similar market patterns map to nearby binary codes regardless of training initialization. The fixed geometric structure also enables the explicit bit‑level splitting required for the s1/s2 hierarchy.
How do I reconstruct the original candlestick values from s1 and s2 tokens?
Reconstruction requires the inverse of the quantization process. While the raw tokens are discrete indices, you must map them back through the quantizer's binary codes and then apply the inverse of the L2‑normalization and scaling used during preprocessing. The KronosPredictor class typically handles this decoding internally during inference to generate price predictions rather than exact reconstruction of input values.
Can I adjust the number of bits allocated to s1 versus s2?
Yes. The codebook_dim parameter in the tokenizer configuration defines the total bit budget (s1 + s2). You can adjust the split by configuring the s1_bits and s2_bits arguments when initializing BSQuantizer or HierarchicalEmbedding. Note that changing these dimensions requires retraining the model, as the embedding layers and dependency‑aware networks are sized to specific token vocabulary ranges (2^s1 and 2^s2).
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →