Kronos Quantization Strategy for K-line Data: A Hybrid Hierarchical-Binary Approach
Kronos employs a hybrid hierarchical-binary quantization strategy that converts continuous OHLCV (Open, High, Low, Close, Volume) time series into discrete tokens through linear projection, binary spherical quantization, and two-stage token splitting.
The shiyu-coder/Kronos repository implements a specialized tokenizer designed specifically for financial K-line data. Understanding its quantization strategy is essential for developers building time series forecasting models, as it transforms raw multivariate vectors into a hierarchical discrete representation that feeds directly into the autoregressive transformer architecture.
The Three-Stage Quantization Pipeline
Kronos processes raw K-line data through a sequential pipeline that reduces high-dimensional continuous values into compact discrete codes. According to the source code in model/kronos.py, the KronosTokenizer class orchestrates this transformation across three distinct stages.
Stage 1: Linear Projection
The pipeline begins with a learnable linear projection layer (self.embed) that maps the raw input dimension d_in to the model dimension d_model. This step standardizes the multivariate OHLCV vectors (typically 6 channels including open, high, low, close, volume, and amount) into a consistent latent space before quantization occurs.
Stage 2: Binary Spherical Quantization
The core quantization logic resides in the BSQuantizer module defined in model/module.py (lines 25-55). This component implements Binary Spherical Quantization through the BinarySphericalQuantizer class:
- Incoming vectors are normalized to unit length on a hypersphere
- The system quantizes continuous values into binary codes consisting of
±1bits - A secondary projection (
self.quant_embed) maps hidden states to the codebook dimension (s1_bits + s2_bits)
This approach effectively compresses the continuous K-line features into a compact binary representation while preserving directional relationships in the high-dimensional space.
Stage 3: Hierarchical Token Splitting
Following quantization, the binary code undergoes a two-stage token split to create hierarchical discrete tokens:
s1_bits(pre-token): Coarse-grained token representing high-level patternss2_bits(post-token): Fine-grained token capturing detailed variations
The HierarchicalEmbedding class (found in model/module.py lines 100-144) processes these separately, embedding quantized_pre and the full quantized representation through distinct linear layers (post_quant_embed_pre and post_quant_embed). This split allows the decoder-only transformer to attend to both coarse and fine temporal structures simultaneously.
Source Code Implementation Details
The quantization strategy is implemented across two primary files in the repository:
model/kronos.py(lines 13-33): Contains theKronosTokenizerclass initialization, which wires together the projection layers and instantiates theBSQuantizermodel/module.py(lines 25-55): Houses theBSQuantizerandBinarySphericalQuantizerclasses that execute the binary spherical codebook logic and bit-to-index conversionmodel/module.py(lines 100-144): ImplementsHierarchicalEmbeddingfor handling the two-stage token embedding (pre and post bits)
Practical Example: Tokenizing OHLCV Data
You can interact with this quantization pipeline directly using the KronosTokenizer class:
import torch
from model import KronosTokenizer
# Example: a batch of 10 daily K-lines, each with 6 channels (open, high, low, close, volume, amount)
batch_size, seq_len, dim = 10, 400, 6
raw_kline = torch.randn(batch_size, seq_len, dim) # replace with real market data
# Load the pre-trained tokenizer
tokenizer = KronosTokenizer.from_pretrained("NeoQuasar/Kronos-Tokenizer-base")
# Encode → get hierarchical token indices (pre-token + post-token)
token_indices = tokenizer.encode(raw_kline) # shape: (batch, seq_len) of ints
# Full forward pass including reconstruction
(z_pre, z), loss, quantized, indices = tokenizer(raw_kline)
print("Pre-token shape :", z_pre.shape) # (batch, seq_len, dim)
print("Full-token shape:", z.shape) # (batch, seq_len, dim)
print("Quantization loss:", loss.item())
This example demonstrates how raw K-line tensors flow through the self.embed projection, into the BSQuantizer, and emerge as hierarchical tokens ready for transformer processing.
Core Components of the Quantization Strategy
Understanding the specific layers involved helps when customizing or debugging the pipeline:
self.embed: Linear projection from raw input dimension to model dimensionself.quant_embed: Projects hidden states to the binary codebook dimension (s1_bits + s2_bits)BSQuantizer: Performs vector normalization and binary spherical quantization, returning binary bits and quantization lossquantized_pre/quantized: Represent the split binary codes (pre-token and full token) that feed into hierarchical embeddingsHierarchicalEmbedding: Converts the two-stage token IDs into continuous embeddings for the autoregressive transformer decoder
Summary
- Kronos implements a hybrid hierarchical-binary quantization strategy specifically designed for OHLCV time series data
- The pipeline uses Binary Spherical Quantization (
BSQuantizer) to convert continuous vectors into±1binary codes - Two-stage token splitting creates coarse (
s1_bits) and fine (s2_bits) hierarchical tokens for multi-scale pattern recognition - Implementation spans
model/kronos.py(tokenizer orchestration) andmodel/module.py(quantization logic and hierarchical embedding) - The
KronosTokenizerclass provides a unified interface for encoding raw K-line data into discrete transformer-ready tokens
Frequently Asked Questions
What is Binary Spherical Quantization in Kronos?
Binary Spherical Quantization is a vector quantization technique implemented in the BSQuantizer class that normalizes input vectors onto a unit hypersphere before mapping them to discrete binary codes (±1 values). This preserves the angular relationships between K-line feature vectors while achieving high compression ratios suitable for discrete token transformers.
Why does Kronos use a two-stage token split?
The two-stage split into s1_bits (pre-token) and s2_bits (post-token) enables hierarchical representation learning. The coarse pre-token captures high-level market trends while the fine-grained post-token preserves detailed price movements, allowing the transformer to attend to patterns at multiple temporal resolutions simultaneously.
How does the HierarchicalEmbedding layer process quantized tokens?
The HierarchicalEmbedding class takes the split binary representations and maps them back to the model dimension through separate linear layers—specifically post_quant_embed_pre for the coarse tokens and post_quant_embed for the full representation. This generates continuous embeddings that the decoder-only transformer can process autoregressively.
Where is the quantization loss calculated?
The quantization loss is computed inside the BSQuantizer forward pass (as seen in the usage pattern (z_pre, z), loss, quantized, indices = tokenizer(raw_kline)). This loss typically measures the commitment cost between the continuous projection and its quantized binary representation, aiding in stable training of the vector quantization layers.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →