Kronos Quantization Strategy for K-line Data: A Hybrid Hierarchical-Binary Approach

Kronos employs a hybrid hierarchical-binary quantization strategy that converts continuous OHLCV (Open, High, Low, Close, Volume) time series into discrete tokens through linear projection, binary spherical quantization, and two-stage token splitting.

The shiyu-coder/Kronos repository implements a specialized tokenizer designed specifically for financial K-line data. Understanding its quantization strategy is essential for developers building time series forecasting models, as it transforms raw multivariate vectors into a hierarchical discrete representation that feeds directly into the autoregressive transformer architecture.

The Three-Stage Quantization Pipeline

Kronos processes raw K-line data through a sequential pipeline that reduces high-dimensional continuous values into compact discrete codes. According to the source code in model/kronos.py, the KronosTokenizer class orchestrates this transformation across three distinct stages.

Stage 1: Linear Projection

The pipeline begins with a learnable linear projection layer (self.embed) that maps the raw input dimension d_in to the model dimension d_model. This step standardizes the multivariate OHLCV vectors (typically 6 channels including open, high, low, close, volume, and amount) into a consistent latent space before quantization occurs.

Stage 2: Binary Spherical Quantization

The core quantization logic resides in the BSQuantizer module defined in model/module.py (lines 25-55). This component implements Binary Spherical Quantization through the BinarySphericalQuantizer class:

  • Incoming vectors are normalized to unit length on a hypersphere
  • The system quantizes continuous values into binary codes consisting of ±1 bits
  • A secondary projection (self.quant_embed) maps hidden states to the codebook dimension (s1_bits + s2_bits)

This approach effectively compresses the continuous K-line features into a compact binary representation while preserving directional relationships in the high-dimensional space.

Stage 3: Hierarchical Token Splitting

Following quantization, the binary code undergoes a two-stage token split to create hierarchical discrete tokens:

  • s1_bits (pre-token): Coarse-grained token representing high-level patterns
  • s2_bits (post-token): Fine-grained token capturing detailed variations

The HierarchicalEmbedding class (found in model/module.py lines 100-144) processes these separately, embedding quantized_pre and the full quantized representation through distinct linear layers (post_quant_embed_pre and post_quant_embed). This split allows the decoder-only transformer to attend to both coarse and fine temporal structures simultaneously.

Source Code Implementation Details

The quantization strategy is implemented across two primary files in the repository:

  • model/kronos.py (lines 13-33): Contains the KronosTokenizer class initialization, which wires together the projection layers and instantiates the BSQuantizer
  • model/module.py (lines 25-55): Houses the BSQuantizer and BinarySphericalQuantizer classes that execute the binary spherical codebook logic and bit-to-index conversion
  • model/module.py (lines 100-144): Implements HierarchicalEmbedding for handling the two-stage token embedding (pre and post bits)

Practical Example: Tokenizing OHLCV Data

You can interact with this quantization pipeline directly using the KronosTokenizer class:

import torch
from model import KronosTokenizer

# Example: a batch of 10 daily K-lines, each with 6 channels (open, high, low, close, volume, amount)

batch_size, seq_len, dim = 10, 400, 6
raw_kline = torch.randn(batch_size, seq_len, dim)   # replace with real market data

# Load the pre-trained tokenizer

tokenizer = KronosTokenizer.from_pretrained("NeoQuasar/Kronos-Tokenizer-base")

# Encode → get hierarchical token indices (pre-token + post-token)

token_indices = tokenizer.encode(raw_kline)   # shape: (batch, seq_len) of ints

# Full forward pass including reconstruction

(z_pre, z), loss, quantized, indices = tokenizer(raw_kline)

print("Pre-token shape :", z_pre.shape)       # (batch, seq_len, dim)

print("Full-token shape:", z.shape)          # (batch, seq_len, dim)

print("Quantization loss:", loss.item())

This example demonstrates how raw K-line tensors flow through the self.embed projection, into the BSQuantizer, and emerge as hierarchical tokens ready for transformer processing.

Core Components of the Quantization Strategy

Understanding the specific layers involved helps when customizing or debugging the pipeline:

  • self.embed: Linear projection from raw input dimension to model dimension
  • self.quant_embed: Projects hidden states to the binary codebook dimension (s1_bits + s2_bits)
  • BSQuantizer: Performs vector normalization and binary spherical quantization, returning binary bits and quantization loss
  • quantized_pre / quantized: Represent the split binary codes (pre-token and full token) that feed into hierarchical embeddings
  • HierarchicalEmbedding: Converts the two-stage token IDs into continuous embeddings for the autoregressive transformer decoder

Summary

  • Kronos implements a hybrid hierarchical-binary quantization strategy specifically designed for OHLCV time series data
  • The pipeline uses Binary Spherical Quantization (BSQuantizer) to convert continuous vectors into ±1 binary codes
  • Two-stage token splitting creates coarse (s1_bits) and fine (s2_bits) hierarchical tokens for multi-scale pattern recognition
  • Implementation spans model/kronos.py (tokenizer orchestration) and model/module.py (quantization logic and hierarchical embedding)
  • The KronosTokenizer class provides a unified interface for encoding raw K-line data into discrete transformer-ready tokens

Frequently Asked Questions

What is Binary Spherical Quantization in Kronos?

Binary Spherical Quantization is a vector quantization technique implemented in the BSQuantizer class that normalizes input vectors onto a unit hypersphere before mapping them to discrete binary codes (±1 values). This preserves the angular relationships between K-line feature vectors while achieving high compression ratios suitable for discrete token transformers.

Why does Kronos use a two-stage token split?

The two-stage split into s1_bits (pre-token) and s2_bits (post-token) enables hierarchical representation learning. The coarse pre-token captures high-level market trends while the fine-grained post-token preserves detailed price movements, allowing the transformer to attend to patterns at multiple temporal resolutions simultaneously.

How does the HierarchicalEmbedding layer process quantized tokens?

The HierarchicalEmbedding class takes the split binary representations and maps them back to the model dimension through separate linear layers—specifically post_quant_embed_pre for the coarse tokens and post_quant_embed for the full representation. This generates continuous embeddings that the decoder-only transformer can process autoregressively.

Where is the quantization loss calculated?

The quantization loss is computed inside the BSQuantizer forward pass (as seen in the usage pattern (z_pre, z), loss, quantized, indices = tokenizer(raw_kline)). This loss typically measures the commitment cost between the continuous projection and its quantized binary representation, aiding in stable training of the vector quantization layers.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →