How Heretic Calculates Refusal Directions: A Deep Dive into the Source Code
Heretic calculates refusal directions by computing the L2-normalized difference between mean residual activations of "bad" prompts (which trigger refusals) and "good" prompts (which yield helpful responses), optionally orthogonalizing the result against the good direction to isolate refusal-specific features.
The p-e-w/heretic repository implements a technique called "abliteration" to modify language model behavior by identifying and neutralizing internal activation patterns that correlate with refusal. Understanding how Heretic calculates refusal directions reveals the statistical foundation of this alignment intervention.
Collecting Residual Activations from Contrastive Prompts
Heretic begins by extracting internal model states for two distinct prompt distributions. In src/heretic/main.py, the pipeline loads batches of good prompts (harmless queries that normally receive compliant responses) and bad prompts (harmful queries that typically trigger safety refusals).
The model computes per-layer residual tensors for both sets:
good_residuals = model.get_residuals_batched(good_prompts)
bad_residuals = model.get_residuals_batched(bad_prompts)
These calls return tensors capturing the hidden state residuals across all transformer layers for each prompt in the batch (see src/heretic/main.py lines 22-25).
Computing Mean Activations Per Layer
To isolate the statistical signature of refusal behavior, Heretic averages the residuals across the prompt dimension. This yields a single representative vector per layer for each prompt category:
good_means = good_residuals.mean(dim=0) # shape: (layers, hidden_dim)
bad_means = bad_residuals.mean(dim=0)
The resulting good_means and bad_means tensors have shape (num_layers, hidden_dim) and represent the centroid of activation space for helpful versus refused responses respectively (see src/heretic/main.py lines 27-29).
Deriving the Initial Refusal Direction
The raw refusal direction emerges from the vector difference between bad and good mean activations. Heretic applies L2 normalization to produce a unit vector pointing toward the refusal region of the embedding space:
refusal_directions = F.normalize(bad_means - good_means, p=2, dim=1)
This operation (see src/heretic/main.py lines 30-31) creates a tensor of shape (layers, hidden_dim) where each row represents the primary direction of variance that distinguishes refused from compliant responses at that specific layer.
Orthogonalizing Against the Good Direction
When the orthogonalize_direction configuration flag is enabled, Heretic refines the refusal direction using a technique inspired by projected abliteration. This removes any component of the refusal vector that aligns with the "good" activation pattern, ensuring the modification targets refusal-specific features rather than general response characteristics.
The implementation in src/heretic/main.py (lines 33-41) projects the raw refusal direction onto the subspace orthogonal to the good mean direction:
good_directions = F.normalize(good_means, p=2, dim=1)
projection_vector = torch.sum(refusal_directions * good_directions, dim=1)
refusal_directions = (
refusal_directions - projection_vector.unsqueeze(1) * good_directions
)
refusal_directions = F.normalize(refusal_directions, p=2, dim=1)
This Gram-Schmidt-like process ensures that the final refusal direction is statistically independent of normal helpful response patterns, reducing the risk of degrading the model's general capabilities during abliteration.
Applying Directions in Model Abliteration
Once calculated, the refusal_directions tensor is passed to the model's abliterate method (defined in src/heretic/model.py lines 386-393). During this phase, Heretic nudges the model's weights opposite to the refusal directions, effectively erasing the internal representations that drive safety refusals. The directions are also cached within the model instance for evaluation and iterative refinement.
Summary
- Residual Extraction: Heretic captures layer-wise activations for contrastive prompt sets using
get_residuals_batched(). - Mean Centering: The system computes per-layer centroids for good and bad prompt distributions to identify activation differences.
- Vector Normalization: Refusal directions are calculated as L2-normalized differences between bad and good means.
- Orthogonal Projection: Optional orthogonalization removes components aligned with good responses, following the projected-abliteration method.
- Weight Intervention: The resulting (layers, hidden_dim) tensor directs the
abliterate()method to modify model weights and reduce refusal behavior.
Frequently Asked Questions
What constitutes "good" and "bad" prompts in Heretic's calculation?
Good prompts are harmless queries that typically elicit helpful, compliant responses from the model, while bad prompts are adversarial or harmful inputs designed to trigger safety refusals. Heretic uses these contrastive sets to statistically isolate the internal activation patterns that correlate with refusal behavior versus normal response generation.
Why does Heretic use L2 normalization when calculating refusal directions?
L2 normalization converts the raw vector difference into a unit vector, ensuring that the refusal direction captures only the orientation of the activation difference rather than its magnitude. This standardization allows consistent application across layers with varying activation scales and prevents the abliteration process from being dominated by high-magnitude outliers in specific layers.
What is the purpose of orthogonalizing refusal directions against good directions?
Orthogonalization removes components of the refusal vector that align with normal helpful responses, ensuring that the abliteration target specifically disrupts refusal mechanisms rather than general model capabilities. This technique, implemented when orthogonalize_direction is enabled, follows the projected-abliteration approach to preserve model performance on benign tasks while eliminating unwanted refusals.
How does Heretic store and apply the calculated refusal directions?
After calculation, Heretic stores the refusal directions tensor (shape: layers × hidden_dim) within the model instance and passes it to the abliterate() method in src/heretic/model.py. This method uses the directions to guide weight modifications that push the model's internal representations away from refusal-associated regions of the activation space.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →