How Heretic's Winsorization Handles Massive Activations in Language Models
Heretic's winsorization clamps extreme residual values by computing a per-prompt, per-layer quantile threshold and symmetrically capping activations to that range, preventing numerical instability in downstream calculations.
Heretic, an open-source tool from the p-e-w/heretic repository, extracts residual vectors from language model hidden states to analyze model behavior. When working with low-precision models or prompts that trigger extreme activation spikes, Heretic's winsorization provides a critical safeguard by truncating outlier values before they propagate to KL-divergence or log-probability computations.
When Heretic's Winsorization Activates
Winsorization occurs at a specific point in the residual extraction pipeline and only under configurable conditions.
Timing in the Processing Pipeline
After the hidden states of each layer are stacked into a tensor of shape (prompt, layer, component), the residuals are cast to float32 for numeric stability. Winsorization is applied immediately after this casting operation, ensuring that subsequent calculations operate on clamped values rather than raw activations.
Configuration Threshold
The feature is controlled by the winsorization_quantile parameter defined in src/heretic/config.py. Winsorization only executes if the user-supplied value lies in the interval [0, 1). The default value of 1.0 explicitly disables the step, ensuring backward compatibility and zero overhead when the feature is not needed.
How Winsorization Computes Quantile Thresholds
The implementation in src/heretic/model.py follows a three-step statistical process to identify and neutralize extreme values.
Absolute Value Calculation
The algorithm first computes the absolute values of all residual components. This creates a non-negative distribution where extreme magnitudes can be statistically identified regardless of their sign.
Per-Slice Quantile Computation
For every (prompt, layer) slice, the quantile of these absolute values is computed along the component dimension (dim=2). This operation yields a threshold tensor of shape (prompt, layer, 1), where each entry represents the activation magnitude at the user-specified quantile for that specific prompt and layer combination.
Symmetric Clamping
The original residuals are then clamped symmetrically to the range [-threshold, threshold]. Any component whose absolute value exceeds the threshold is set to exactly the threshold value, preserving the sign but capping the magnitude. The clamped tensor is returned for downstream processing; if winsorization is disabled, the untouched residuals flow through unchanged.
Configuration and Implementation Details
The winsorization pipeline spans two critical source files:
src/heretic/config.py: Defines thewinsorization_quantilefield, its default value of1.0, and documentation explaining that values in[0, 1)enable the feature while1.0disables it.src/heretic/model.py: Contains the core logic for casting tofloat32, computing quantile thresholds alongdim=2, and applying the symmetric clamp operation.
Practical Examples
Enabling Winsorization via Configuration
Create a config.toml file to set the 95th percentile as your clamping threshold:
[settings]
winsorization_quantile = 0.95
This configuration tames large activation spikes by capping residuals at the 95th percentile of their absolute magnitude.
Loading and Using Winsorized Residuals
from heretic import Heretic, Settings
# Load the model with the custom setting
settings = Settings.from_toml("config.toml")
heretic = Heretic(settings)
# Obtain residuals – winsorization is applied automatically
prompts = ["Explain quantum entanglement.", "Give a short poem."]
residuals = heretic.get_residuals(prompts) # shape: (2, L, C)
print(residuals.shape) # → (2, num_layers, component_dim)
# The residuals have been symmetrically clamped according to the 0.95-quantile.
Disabling Winsorization
To return raw residuals without clamping, set the quantile to 1.0 or greater:
settings.winsorization_quantile = 1.0 # Disables the winsorization step
heretic = Heretic(settings)
raw_residuals = heretic.get_residuals(prompts)
Summary
- Heretic's winsorization clamps residual vectors to prevent massive activations from destabilizing downstream statistical calculations.
- The feature activates when
winsorization_quantileis set in the interval [0, 1), with a default of1.0disabling the behavior. - Implementation in
src/heretic/model.pycomputes per-prompt, per-layer quantile thresholds along the component dimension and applies symmetric clamping to[-threshold, threshold]. - This safeguard is particularly critical when analyzing low-precision models (e.g.,
bfloat16) where extreme activation spikes are common.
Frequently Asked Questions
What is the default winsorization behavior in Heretic?
By default, Heretic disables winsorization by setting winsorization_quantile = 1.0 in src/heretic/config.py. This ensures zero overhead and backward compatibility, returning raw residual vectors without any clamping unless the user explicitly configures a quantile in the range [0, 1).
How does winsorization affect model performance and accuracy?
Winsorization specifically targets numerical stability rather than model accuracy. By clamping extreme residual components to a quantile-based threshold, it prevents massive activation values from skewing KL-divergence or log-probability calculations. The process operates on extracted residuals after forward passes, so it does not modify the underlying language model weights or inference behavior.
Can I apply different winsorization quantiles to different layers?
The current implementation in src/heretic/model.py computes quantiles along the component dimension (dim=2) for each (prompt, layer) pair using a single global winsorization_quantile value. While this applies the same quantile threshold across all layers, the actual threshold value varies per-layer based on the distribution of activations in that specific layer. Per-layer quantile configuration is not currently supported in the configuration schema.
Why is float32 casting necessary before winsorization?
The residuals are cast to float32 before quantile computation to ensure numeric stability when calculating statistics across potentially millions of activation values. Low-precision formats like bfloat16 or float16 can suffer from overflow or precision loss during quantile calculations, especially when dealing with the extreme values that winsorization is designed to handle. The float32 cast prevents these numerical artifacts from affecting the threshold computation.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →