BSQuantizer Entropy Penalty Terms (γ₀, γ, ζ) in Kronos Codebook Learning

The BSQuantizer entropy penalty terms γ₀, γ, and ζ control binary codebook diversity by regularizing per-sample entropy, codebook entropy, and overall regularization strength respectively.

The BSQuantizer class in the Kronos repository implements vector quantization using a binary spherical codebook defined in model/module.py. During training, three entropy penalty hyperparameters—γ₀ (gamma0), γ (gamma), and ζ (zeta)—govern how binary codes are distributed across samples and within the global codebook.

How Entropy Regularization Works in BinarySphericalQuantizer

Inside BinarySphericalQuantizer.forward, the quantizer computes a composite entropy penalty that balances local usage patterns against global codebook coverage. According to the implementation in model/module.py at lines 109–112 and line 126, the loss construction follows this pattern:

entropy_penalty = self.gamma0 * persample_entropy - self.gamma * cb_entropy

# ...

commit_loss + self.zeta * entropy_penalty / self.inv_temperature

The final bsq_loss returned to the trainer combines the commitment loss with this scaled entropy penalty, directly influencing gradient updates during backpropagation.

Gamma0 (γ₀): Per-Sample Entropy Weight

The gamma0 parameter scales the per-sample entropy (persample_entropy). When you increase γ₀, the optimizer receives stronger pressure to maximize entropy across individual samples. This encourages each bit position in the binary code to activate uniformly across the batch, preventing the model from relying on a sparse subset of bits for every input vector.

Gamma (γ): Codebook Entropy Weight

The gamma parameter weights the codebook entropy (cb_entropy). Because this term is subtracted from the loss calculation, larger values of γ reduce the total loss when the codebook distribution exhibits high entropy. This effectively rewards the model for populating a larger portion of the available binary space, ensuring the codebook remains diverse and expressive rather than collapsing to a few dominant codes.

Zeta (ζ): Global Regularization Scaling

The zeta parameter acts as a master scaling factor for the combined entropy penalty. Applied at line 126, ζ determines how strongly the entropy regularization influences the overall training objective relative to the commitment loss. Setting ζ to 0.0 disables entropy regularization entirely, while values greater than 1.0 amplify the influence of both γ₀ and γ on the learned binary representations.

Practical Configuration Examples

The following examples demonstrate how to instantiate BSQuantizer with different entropy penalty configurations to achieve specific training behaviors:

import torch
from model.module import BSQuantizer

# Dummy input tensor (batch, seq_len, embed_dim)

z = torch.randn(4, 10, 64)

# Standard configuration with moderate entropy regularization

quant = BSQuantizer(
    s1_bits=32,
    s2_bits=32,
    beta=0.25,
    gamma0=0.1,    # Per-sample entropy weight

    gamma=0.05,    # Codebook entropy weight

    zeta=1.0,      # Global scaling

    group_size=8,
)
bsq_loss, q_vec, indices = quant(z)

# Strong per-sample regularization (higher gamma0)

quant_strong_ps = BSQuantizer(
    s1_bits=32, s2_bits=32,
    beta=0.25,
    gamma0=0.5,    # Increased per-sample pressure

    gamma=0.05,
    zeta=1.0,
    group_size=8,
)

# Disabling entropy regularization entirely (zeta = 0)

quant_no_entropy = BSQuantizer(
    s1_bits=32, s2_bits=32,
    beta=0.25,
    gamma0=0.1,
    gamma=0.05,
    zeta=0.0,      # Entropy penalty ignored

    group_size=8,
)

Changing these hyperparameters directly alters the loss landscape, forcing the binary codebook to adapt its structure during the forward pass defined in model/module.py.

Impact on Training Dynamics

The interplay between γ₀ and γ creates a tension between local uniformity and global diversity. High γ₀ ensures no single bit dominates individual samples, while high γ prevents the codebook from collapsing to a narrow subset of binary codes. The ζ parameter provides coarse control over this trade-off without requiring individual adjustment of the entropy components. In practice, practitioners often tune γ₀ and γ separately to diagnose codebook collapse, then adjust ζ to balance reconstruction quality against entropy regularization.

Summary

  • γ₀ (gamma0) controls per-sample bit usage uniformity by weighting per-sample entropy in model/module.py lines 109–112
  • γ (gamma) encourages codebook diversity by subtracting a weighted codebook entropy term from the loss
  • ζ (zeta) scales the total entropy penalty before adding it to the commitment loss at line 126
  • Setting ζ to 0.0 effectively disables all entropy regularization, reducing training to commit-loss only
  • All parameters are implemented within BinarySphericalQuantizer.forward and consumed by BSQuantizer during model training

Frequently Asked Questions

What is the difference between gamma0 and gamma in BSQuantizer?

Gamma0 (γ₀) regularizes the entropy of individual sample codes, encouraging each bit to be active roughly half the time across a batch. Gamma (γ) regularizes the entropy of the entire codebook distribution, encouraging the model to utilize many different binary codes rather than collapsing to a few popular ones. Gamma0 affects local sample statistics while gamma affects global codebook utilization.

How do I completely disable entropy regularization during training?

Set the zeta parameter to 0.0 when instantiating BSQuantizer. Because ζ scales the combined entropy penalty term at line 126 of model/module.py, a value of zero nullifies both the γ₀ and γ components regardless of their individual values, reducing the loss to only the commitment loss weighted by beta.

Why is the gamma term subtracted in the entropy penalty calculation?

The gamma term is subtracted (- self.gamma * cb_entropy) because maximizing codebook entropy reduces the loss. Unlike the per-sample entropy term which is added to increase loss when samples have low entropy, the subtraction means that higher codebook entropy directly lowers the total bsq_loss. This inversion aligns the gradient direction to reward diverse codebook usage.

Where are the entropy penalty terms initialized in the Kronos codebase?

The parameters γ₀, γ, and ζ are initialized as constructor arguments in BSQuantizer and passed through to BinarySphericalQuantizer.__init__ in model/module.py. They are stored as instance attributes (self.gamma0, self.gamma, self.zeta) and accessed during the forward pass at lines 109–112 and line 126 to compute the final entropy penalty.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →