# How Kronos Tokenizes Financial Candlestick Data into s1 and s2 Tokens

> Learn how Kronos tokenizes financial candlestick data into s1 and s2 tokens. Discover the Binary Spherical Quantizer and its two-level hierarchical token stream for efficient data representation.

- Repository: [ShiYu/Kronos](https://github.com/shiyu-coder/Kronos)
- Tags: internals
- Published: 2026-04-10

---

**Kronos converts raw OHLCV candlestick data into a two-level hierarchical token stream using a Binary Spherical Quantizer that splits binary representations into high-order s1 (pre‑token) and low-order s2 (post‑token) indices.**

Kronos is an open‑source financial time‑series model that represents market data as discrete tokens rather than continuous values. The tokenization process transforms normalized price‑volume tensors into compact binary codes that are split into **s1** and **s2** tokens, enabling the model to capture market structure at both coarse and fine granularities.

## The Three-Stage Tokenization Pipeline

The tokenization flow implemented in `shiyu‑coder/Kronos` follows a clear pipeline from raw dataframe to paired token indices. According to the source code in [`model/kronos.py`](https://github.com/shiyu-coder/Kronos/blob/main/model/kronos.py) and [`model/module.py`](https://github.com/shiyu-coder/Kronos/blob/main/model/module.py), the process involves normalization, binary‑spherical quantization, and hierarchical splitting.

### Stage 1: Normalization and Tensor Preparation

Before quantization, the `KronosPredictor` normalizes the raw OHLCV (open‑high‑low‑close‑volume) dataframe and constructs a 3‑D tensor `x` with shape `[batch, seq_len, feat]`. This normalized tensor serves as the continuous input to the tokenizer, ensuring that price and volume scales do not bias the subsequent binary encoding.

### Stage 2: Binary-Spherical Quantization

The core quantization logic resides in `BSQuantizer` (Binary Spherical Quantizer) defined in [`model/module.py`](https://github.com/shiyu-coder/Kronos/blob/main/model/module.py). When `KronosTokenizer.encode(..., half=True)` is invoked, the normalized tensor flows through the following steps:

1. **L2 Normalization** – Features are projected onto a unit sphere to remove magnitude variations.
2. **Binary Code Generation** – The quantizer produces a binary code of length `codebook_dim` (equal to `s1 + s2` bits).
3. **Bit‑to‑Index Conversion** – The `bits_to_indices` method interprets the binary vector as an integer using power‑of‑two weighting.

```python

# model/module.py – BSQuantizer.bits_to_indices (lines 45-55 context)

def bits_to_indices(self, bits):
    bits = (bits >= 0).to(torch.long)
    indices = 2 ** torch.arange(
        0,
        bits.shape[-1],
        1,
        dtype=torch.long,
        device=bits.device,
    )
    return (bits * indices).sum(-1)

```

When the `half=True` flag is passed to `encode()`, the full binary code is split into two contiguous segments: the first `s1` bits and the last `s2` bits. Each segment is converted to a token index tensor independently, yielding two separate outputs rather than a single composite ID.

### Stage 3: Hierarchical Token Splitting

For scenarios where tokens are represented as a single composite integer (e.g., when loading pre‑computed token IDs), the `HierarchicalEmbedding` class provides a `split_token` method to decompose the integer back into s1 and s2 components using bitwise operations:

```python

# model/module.py – HierarchicalEmbedding.split_token (lines 17-28)

def split_token(self, token_ids: torch.Tensor, s2_bits: int):
    t = token_ids.long()
    mask = (1 << s2_bits) - 1
    s2_ids = t & mask           # low bits → s₂

    s1_ids = t >> s2_bits       # high bits → s₁

    return s1_ids, s2_ids

```

This utility allows the model to accept either pre‑split token pairs or unified tokens that get separated inside the embedding layer.

## Practical Implementation: Tokenizing OHLCV Data

The following example demonstrates the end‑to‑end tokenization of a candlestick CSV using the public `KronosTokenizer`:

```python
import pandas as pd
import torch
from model.kronos import KronosTokenizer

# Load OHLCV data

df = pd.read_csv("examples/data/XSHG_5min_600977.csv")
price_cols = ["open", "high", "low", "close"]
df = df[price_cols + ["volume"]]

# Initialize tokenizer (s₁=8 bits, s₂=8 bits in base model)

tokenizer = KronosTokenizer.from_pretrained(
    "NeoQuasar/Kronos-Tokenizer-base", revision="v1.0"
)

# Prepare tensor: [batch=1, time, features]

x_tensor = torch.from_numpy(df.values.astype("float32")).unsqueeze(0)

# Encode with half=True to receive separate s1 and s2 tensors

s1_ids, s2_ids = tokenizer.encode(x_tensor, half=True)

print(f"s₁ shape: {s1_ids.shape}")  # torch.Size([1, T])

print(f"s₂ shape: {s2_ids.shape}")  # torch.Size([1, T])

```

The `encode` method in [`model/kronos.py`](https://github.com/shiyu-coder/Kronos/blob/main/model/kronos.py) (lines 42‑53) orchestrates this call, forwarding the tensor to `BSQuantizer.forward` and returning the split indices when `half=True` is specified.

## Understanding s1 and s2 Token Roles

The hierarchical design separates market information into **interdependent abstraction levels**:

- **s1 Tokens (Pre‑tokens)** – Represented by the **high‑order bits** of the binary code, these act as a coarse vocabulary. The model first processes s1 tokens to establish a broad market context or "theme" for the current time step.
- **s2 Tokens (Post‑tokens)** – Represented by the **low‑order bits**, these provide fine‑grained detail. Within the transformer architecture, s2 generation is conditioned on the decoded s1 context through a dependency‑aware layer, ensuring that granular price movements remain grounded in the broader market structure established by s1.

This bit‑level separation allows Kronos to model the conditional distribution `P(s2 | s1)`, effectively learning hierarchical dependencies in financial time series while maintaining a compact codebook of size `2^(s1+s2)`.

## Summary

- **Kronos** tokenizes financial candlestick data via **Binary Spherical Quantization** in [`model/module.py`](https://github.com/shiyu-coder/Kronos/blob/main/model/module.py).
- The `BSQuantizer` generates binary codes of length `s1 + s2` bits via L2‑normalized projection.
- Setting `half=True` in `KronosTokenizer.encode()` splits the binary string into separate **s1** (high‑order) and **s2** (low‑order) token tensors.
- `HierarchicalEmbedding.split_token` can decompose composite token IDs using bitwise masks and shifts.
- **s1** captures coarse market structure while **s2** refines it, enabling conditional generation within the model architecture.

## Frequently Asked Questions

### What do s1 and s2 stand for in Kronos tokenization?

**s1** refers to "pre‑tokens" (high‑order bits) and **s2** refers to "post‑tokens" (low‑order bits). This naming reflects the hierarchical generation order: the model first predicts or attends to s1 tokens to establish context, then generates s2 tokens conditioned on that context, similar to a coarse‑to‑fine quantization strategy.

### Why does Kronos use binary quantization instead of learned vector quantization?

Binary Spherical Quantization maps normalized price movements to vertices of a hypercube inscribed on a unit sphere. This approach provides **deterministic, reversible encoding** without requiring a learned codebook, ensuring that similar market patterns map to nearby binary codes regardless of training initialization. The fixed geometric structure also enables the explicit bit‑level splitting required for the s1/s2 hierarchy.

### How do I reconstruct the original candlestick values from s1 and s2 tokens?

Reconstruction requires the inverse of the quantization process. While the raw tokens are discrete indices, you must map them back through the quantizer's binary codes and then apply the inverse of the L2‑normalization and scaling used during preprocessing. The `KronosPredictor` class typically handles this decoding internally during inference to generate price predictions rather than exact reconstruction of input values.

### Can I adjust the number of bits allocated to s1 versus s2?

Yes. The `codebook_dim` parameter in the tokenizer configuration defines the total bit budget (`s1 + s2`). You can adjust the split by configuring the `s1_bits` and `s2_bits` arguments when initializing `BSQuantizer` or `HierarchicalEmbedding`. Note that changing these dimensions requires retraining the model, as the embedding layers and dependency‑aware networks are sized to specific token vocabulary ranges (`2^s1` and `2^s2`).