What Is the Binary Spherical Quantizer (BSQuantizer) in Kronos and How Does It Compress OHLCV Data?

The Binary Spherical Quantizer (BSQuantizer) is a hybrid tokenization component in Kronos that compresses OHLCV time-series data by L2-normalizing latent vectors, quantizing them into binary codes on the unit hypersphere, and collapsing those bits into compact integer indices.

The Kronos tokenizer uses BSQuantizer as its core compression engine to turn high-dimensional financial data into discrete tokens suitable for transformer models. This component, implemented in model/module.py, wraps the Binary Spherical Quantization algorithm to achieve drastic bitrate reduction while preserving geometric relationships necessary for accurate market reconstruction.

How BSQuantizer Works

The quantizer lives in model/module.py (lines 25–55) and implements a three-stage pipeline that converts continuous latent representations into discrete codes.

Step 1: L2 Normalization

First, the input tensor z—the output of a linear projection following the transformer encoder—is projected onto the unit hypersphere. This normalization ensures all vectors have equal magnitude, forcing the model to encode information solely in directional relationships.

In module.py lines 45–46, the operation is:

z = F.normalize(z, dim=-1)

Step 2: Binary Spherical Quantization

The normalized vectors pass into BinarySphericalQuantizer (instantiated as self.bsq), which performs three operations:

  • Rounding: Each dimension is quantized to ±1 using the quantize function
  • Soft-entropy penalty: Encourages balanced use of the binary codebook during training
  • Commitment loss: Computes β × ‖z − ẑ‖² to keep quantized vectors close to originals

The forward call occurs at line 47:

quantized, bsq_loss, metrics = self.bsq(z, collect_metrics=collect_metrics)

Step 3: Bit-to-Index Conversion

After quantization, the binary tensor is converted to integer indices via the bits_to_indices method (lines 34–44). Each bit-plane is interpreted as a binary digit and summed with powers of two (2**i). When half=True, the codebook splits into separate pre (s1_bits) and post (s2_bits) tensors, yielding two index streams.

The final output returns three objects (line 54):

return bsq_loss, quantized, z_indices

OHLCV Compression Pipeline

BSQuantizer serves as the bottleneck in Kronos’s end-to-end tokenization workflow, reducing raw market data to a fraction of its original size.

From Raw Prices to Discrete Tokens

The compression follows this strict sequence:

  1. Embedding: Raw Open-High-Low-Close-Volume (OHLCV) series undergo linear projection via self.embed to the model dimension
  2. Encoding: Several transformer encoder blocks process the sequence, producing latent representation z
  3. Projection: z maps to a vector of size s1_bits + s2_bits through self.quant_embed—a dimension drastically smaller than the original input
  4. Quantization: BSQuantizer normalizes, binarizes, and packs the vector into integer indices (z_indices)
  5. Storage: Each timestep occupies exactly s1_bits + s2_bits bits, replacing the original 32-bit float matrix

For a typical 400-timestep window, storage drops from 400 × 5 × 32 bits (raw OHLCV floats) to 400 × (s1_bits + s2_bits) bits—often a 10× or greater reduction.

Public API Integration

The KronosTokenizer class in model/kronos.py exposes this compression through two methods:

  • encode (lines 42–59): Runs the pipeline through step 4, returning only the compressed integer indices
  • decode (lines 61–77): Reverses the process, converting indices back to approximate OHLCV reconstructions

Implementation Details

BSQuantizer Class Structure

The implementation wraps the research-grade BinarySphericalQuantizer from the paper “Binary Spherical Quantization” (arXiv:2406.07548). Key methods include:

  • forward: Orchestrates normalization, quantization, and index extraction
  • bits_to_indices: Handles binary-to-integer packing logic
  • quantize: Applies the ±1 rounding operation

Source reference: model/module.py lines 90–104 contain the core forward logic for quantization and loss computation.

Integration with KronosTokenizer

In model/kronos.py lines 70–74, the tokenizer instantiates BSQuantizer:

self.quantizer = BSQuantizer(
    latent_dim=config.model_dim,
    s1_bits=config.s1_bits,
    s2_bits=config.s2_bits,
    beta=config.beta,
)

During the forward pass (lines 89–96), embedded OHLCV flows through the encoder and into the quantizer, producing the compressed token stream used for downstream forecasting.

Code Examples

Encoding OHLCV Windows

The following example demonstrates compressing a 400-timestep OHLCV window into discrete tokens:

import pandas as pd
import torch
from model import KronosTokenizer

# Load raw OHLCV data (shape: [lookback, 5])

df = pd.read_csv("data/XSHG_5min_600977.csv")
lookback = 400
x_df = df.loc[:lookback-1, ["open", "high", "low", "close", "volume"]]

# Convert to tensor [batch, time, features]

x = torch.tensor(x_df.values, dtype=torch.float32).unsqueeze(0)

# Initialize tokenizer

tokenizer = KronosTokenizer.from_pretrained("NeoQuasar/Kronos-Tokenizer-base")

# Compress to integer indices

z_indices = tokenizer.encode(x)
print(f"Compressed shape: {z_indices.shape}")  # [1, 400]

print(f"Bits per timestep: {tokenizer.s1_bits + tokenizer.s2_bits}")

# Reconstruct approximate OHLCV

reconstructed = tokenizer.decode(z_indices)

Half-Precision Mode

For split-codebook scenarios, use the half=True flag to generate separate pre- and post-token indices:


# Encode into two separate index tensors

z_pre, z_post = tokenizer.encode(x, half=True)

# Decode each component

rec_pre = tokenizer.decode(z_pre, half=True, part='pre')
rec_post = tokenizer.decode(z_post, half=True, part='post')

# Full reconstruction combines both parts

reconstructed = tokenizer.decode([z_pre, z_post], half=True)

Summary

  • BSQuantizer implements binary spherical quantization in model/module.py, wrapping the algorithm from arXiv:2406.07548
  • The pipeline L2-normalizes vectors, quantizes to ±1 binary codes, and packs bits into integer indices via bits_to_indices
  • OHLCV compression reduces storage from dense 32-bit floats to s1_bits + s2_bits per timestep—typically achieving 10× or greater size reduction
  • KronosTokenizer.encode returns compressed indices; decode reconstructs the approximate time series
  • The commit loss (β term) ensures quantized representations remain geometrically close to the original latent space

Frequently Asked Questions

How does BSQuantizer differ from standard VQ-VAE quantization?

Standard VQ-VAE uses learned codebook embeddings and nearest-neighbor lookup, which requires storing full embedding vectors. BSQuantizer instead constrains vectors to the unit hypersphere and uses deterministic ±1 rounding, eliminating the need to store a large codebook while maintaining reconstruction fidelity through geometric preservation.

What is the purpose of the half=True mode in the tokenizer?

The half=True mode splits the quantization into two stages (s1_bits and s2_bits), generating separate index tensors for the first and second halves of the binary code. This enables hierarchical tokenization where pre-tokens capture coarse market structure and post-tokens encode fine-grained movements, useful for multi-resolution forecasting models.

Why is L2 normalization performed before quantization?

L2 normalization (implemented via F.normalize(z, dim=-1)) forces all latent vectors onto the unit hypersphere. This constraint removes magnitude information and ensures that quantization depends solely on directional similarity, which is better preserved by binary spherical quantization than by standard Euclidean distance metrics.

Can BSQuantizer handle multivariate time series beyond OHLCV?

Yes. While optimized for financial OHLCV data, the quantizer operates on any tensor of shape [batch, time, features]. The d_in dimension (input features) projects to s1_bits + s2_bits regardless of source, meaning the compression mechanism generalizes to any multivariate time series that fits the Kronos encoder architecture.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →