# What Is the Binary Spherical Quantizer (BSQuantizer) in Kronos and How Does It Compress OHLCV Data?

> Learn how the Binary Spherical Quantizer BSQuantizer in Kronos compresses OHLCV data by L2-normalizing latent vectors and quantizing them into binary codes on the unit hypersphere.

- Repository: [ShiYu/Kronos](https://github.com/shiyu-coder/Kronos)
- Tags: deep-dive
- Published: 2026-04-10

---

**The Binary Spherical Quantizer (BSQuantizer) is a hybrid tokenization component in Kronos that compresses OHLCV time-series data by L2-normalizing latent vectors, quantizing them into binary codes on the unit hypersphere, and collapsing those bits into compact integer indices.**

The Kronos tokenizer uses BSQuantizer as its core compression engine to turn high-dimensional financial data into discrete tokens suitable for transformer models. This component, implemented in [`model/module.py`](https://github.com/shiyu-coder/Kronos/blob/main/model/module.py), wraps the Binary Spherical Quantization algorithm to achieve drastic bitrate reduction while preserving geometric relationships necessary for accurate market reconstruction.

## How BSQuantizer Works

The quantizer lives in [`model/module.py`](https://github.com/shiyu-coder/Kronos/blob/main/model/module.py) (lines 25–55) and implements a three-stage pipeline that converts continuous latent representations into discrete codes.

### Step 1: L2 Normalization

First, the input tensor `z`—the output of a linear projection following the transformer encoder—is projected onto the unit hypersphere. This normalization ensures all vectors have equal magnitude, forcing the model to encode information solely in directional relationships.

In [`module.py`](https://github.com/shiyu-coder/Kronos/blob/main/module.py) lines 45–46, the operation is:

```python
z = F.normalize(z, dim=-1)

```

### Step 2: Binary Spherical Quantization

The normalized vectors pass into `BinarySphericalQuantizer` (instantiated as `self.bsq`), which performs three operations:

- **Rounding**: Each dimension is quantized to **±1** using the `quantize` function
- **Soft-entropy penalty**: Encourages balanced use of the binary codebook during training
- **Commitment loss**: Computes `β × ‖z − ẑ‖²` to keep quantized vectors close to originals

The forward call occurs at line 47:

```python
quantized, bsq_loss, metrics = self.bsq(z, collect_metrics=collect_metrics)

```

### Step 3: Bit-to-Index Conversion

After quantization, the binary tensor is converted to integer indices via the `bits_to_indices` method (lines 34–44). Each bit-plane is interpreted as a binary digit and summed with powers of two (`2**i`). When `half=True`, the codebook splits into separate **pre** (`s1_bits`) and **post** (`s2_bits`) tensors, yielding two index streams.

The final output returns three objects (line 54):

```python
return bsq_loss, quantized, z_indices

```

## OHLCV Compression Pipeline

BSQuantizer serves as the bottleneck in Kronos’s end-to-end tokenization workflow, reducing raw market data to a fraction of its original size.

### From Raw Prices to Discrete Tokens

The compression follows this strict sequence:

1. **Embedding**: Raw Open-High-Low-Close-Volume (OHLCV) series undergo linear projection via `self.embed` to the model dimension
2. **Encoding**: Several transformer encoder blocks process the sequence, producing latent representation `z`
3. **Projection**: `z` maps to a vector of size `s1_bits + s2_bits` through `self.quant_embed`—a dimension drastically smaller than the original input
4. **Quantization**: BSQuantizer normalizes, binarizes, and packs the vector into integer indices (`z_indices`)
5. **Storage**: Each timestep occupies exactly **s1_bits + s2_bits** bits, replacing the original 32-bit float matrix

For a typical 400-timestep window, storage drops from **400 × 5 × 32 bits** (raw OHLCV floats) to **400 × (s1_bits + s2_bits)** bits—often a 10× or greater reduction.

### Public API Integration

The `KronosTokenizer` class in [`model/kronos.py`](https://github.com/shiyu-coder/Kronos/blob/main/model/kronos.py) exposes this compression through two methods:

- **`encode`** (lines 42–59): Runs the pipeline through step 4, returning only the compressed integer indices
- **`decode`** (lines 61–77): Reverses the process, converting indices back to approximate OHLCV reconstructions

## Implementation Details

### BSQuantizer Class Structure

The implementation wraps the research-grade `BinarySphericalQuantizer` from the paper *“Binary Spherical Quantization”* (arXiv:2406.07548). Key methods include:

- `forward`: Orchestrates normalization, quantization, and index extraction
- `bits_to_indices`: Handles binary-to-integer packing logic
- `quantize`: Applies the ±1 rounding operation

Source reference: [`model/module.py`](https://github.com/shiyu-coder/Kronos/blob/main/model/module.py) lines 90–104 contain the core forward logic for quantization and loss computation.

### Integration with KronosTokenizer

In [`model/kronos.py`](https://github.com/shiyu-coder/Kronos/blob/main/model/kronos.py) lines 70–74, the tokenizer instantiates BSQuantizer:

```python
self.quantizer = BSQuantizer(
    latent_dim=config.model_dim,
    s1_bits=config.s1_bits,
    s2_bits=config.s2_bits,
    beta=config.beta,
)

```

During the forward pass (lines 89–96), embedded OHLCV flows through the encoder and into the quantizer, producing the compressed token stream used for downstream forecasting.

## Code Examples

### Encoding OHLCV Windows

The following example demonstrates compressing a 400-timestep OHLCV window into discrete tokens:

```python
import pandas as pd
import torch
from model import KronosTokenizer

# Load raw OHLCV data (shape: [lookback, 5])

df = pd.read_csv("data/XSHG_5min_600977.csv")
lookback = 400
x_df = df.loc[:lookback-1, ["open", "high", "low", "close", "volume"]]

# Convert to tensor [batch, time, features]

x = torch.tensor(x_df.values, dtype=torch.float32).unsqueeze(0)

# Initialize tokenizer

tokenizer = KronosTokenizer.from_pretrained("NeoQuasar/Kronos-Tokenizer-base")

# Compress to integer indices

z_indices = tokenizer.encode(x)
print(f"Compressed shape: {z_indices.shape}")  # [1, 400]

print(f"Bits per timestep: {tokenizer.s1_bits + tokenizer.s2_bits}")

# Reconstruct approximate OHLCV

reconstructed = tokenizer.decode(z_indices)

```

### Half-Precision Mode

For split-codebook scenarios, use the `half=True` flag to generate separate pre- and post-token indices:

```python

# Encode into two separate index tensors

z_pre, z_post = tokenizer.encode(x, half=True)

# Decode each component

rec_pre = tokenizer.decode(z_pre, half=True, part='pre')
rec_post = tokenizer.decode(z_post, half=True, part='post')

# Full reconstruction combines both parts

reconstructed = tokenizer.decode([z_pre, z_post], half=True)

```

## Summary

- **BSQuantizer** implements binary spherical quantization in [`model/module.py`](https://github.com/shiyu-coder/Kronos/blob/main/model/module.py), wrapping the algorithm from arXiv:2406.07548
- The pipeline **L2-normalizes** vectors, quantizes to **±1 binary codes**, and packs bits into **integer indices** via `bits_to_indices`
- OHLCV compression reduces storage from dense 32-bit floats to **s1_bits + s2_bits** per timestep—typically achieving **10× or greater** size reduction
- **`KronosTokenizer.encode`** returns compressed indices; **`decode`** reconstructs the approximate time series
- The commit loss (`β` term) ensures quantized representations remain geometrically close to the original latent space

## Frequently Asked Questions

### How does BSQuantizer differ from standard VQ-VAE quantization?

Standard VQ-VAE uses learned codebook embeddings and nearest-neighbor lookup, which requires storing full embedding vectors. BSQuantizer instead constrains vectors to the unit hypersphere and uses deterministic **±1 rounding**, eliminating the need to store a large codebook while maintaining reconstruction fidelity through geometric preservation.

### What is the purpose of the `half=True` mode in the tokenizer?

The `half=True` mode splits the quantization into two stages (`s1_bits` and `s2_bits`), generating separate index tensors for the first and second halves of the binary code. This enables hierarchical tokenization where pre-tokens capture coarse market structure and post-tokens encode fine-grained movements, useful for multi-resolution forecasting models.

### Why is L2 normalization performed before quantization?

L2 normalization (implemented via `F.normalize(z, dim=-1)`) forces all latent vectors onto the unit hypersphere. This constraint removes magnitude information and ensures that quantization depends solely on **directional similarity**, which is better preserved by binary spherical quantization than by standard Euclidean distance metrics.

### Can BSQuantizer handle multivariate time series beyond OHLCV?

Yes. While optimized for financial OHLCV data, the quantizer operates on any tensor of shape `[batch, time, features]`. The `d_in` dimension (input features) projects to `s1_bits + s2_bits` regardless of source, meaning the compression mechanism generalizes to any multivariate time series that fits the Kronos encoder architecture.