Benefits of Using orthogonalize_direction in Heretic for Targeted Model Abliteration

Enabling orthogonalize_direction in Heretic removes the component of refusal directions that overlaps with "good" model behavior, preserving helpful capabilities while eliminating harmful refusals through projected abliteration.

Heretic is an open-source tool for model alignment that computes refusal directions—vectors pointing from a model's acceptable activation space toward refusal-generating regions. When you enable the orthogonalize_direction setting, the algorithm applies projected abliteration by stripping away the parallel component of refusal directions, keeping only the orthogonal part. This precision targeting makes safety tuning less destructive and more efficient according to the p-e-w/heretic source code.

What Does orthogonalize_direction Do in Heretic?

In the standard abliteration workflow, Heretic identifies directions in activation space that trigger refusals. However, these refusal directions often contain components that overlap with the model's "good" directions—vectors representing helpful, harmless responses.

When orthogonalize_direction is enabled, Heretic performs a mathematical projection that removes the component of the refusal direction that lies along the good-direction vector. The result is a purified refusal direction that exclusively represents harmful refusal behavior, without encroaching on the model's useful knowledge base.

Implementation in the Heretic Source Code

The orthogonalization logic is implemented in src/heretic/main.py, where the algorithm adjusts refusal directions before they are used in abliteration:

if settings.orthogonalize_direction:
    # Adjust the refusal directions so that only the component that is

    # orthogonal to the good direction is subtracted during abliteration.

    good_directions = F.normalize(good_means, p=2, dim=1)
    projection_vector = torch.sum(refusal_directions * good_directions, dim=1)
    refusal_directions = (
        refusal_directions - projection_vector.unsqueeze(1) * good_directions
    )
    refusal_directions = F.normalize(refusal_directions, p=2, dim=1)

This code computes the projection of refusal directions onto good directions, subtracts this parallel component, and renormalizes the result. The setting itself is defined in src/heretic/config.py (lines 179-184) within the Settings class schema with the description: "Whether to adjust the refusal directions so that only the component that is orthogonal to the good direction is subtracted during abliteration."

Key Benefits of Using orthogonalize_direction

More Targeted Abliteration

By discarding the parallel component, orthogonalize_direction ensures Heretic only suppresses the specific neural pathways that drive refusals. This prevents the "collateral damage" that can occur when abliteration accidentally modifies weights responsible for legitimate reasoning capabilities.

Preservation of Model Capabilities

The orthogonal component specifically avoids interfering with the model's ability to answer harmless prompts. When the refusal direction is purified through orthogonalization, the resulting model maintains its performance on standard benchmarks and helpful queries, resulting in fewer unintended capability degradations.

Improved Safety-Utility Trade-off

Experimental results demonstrate that using orthogonalized directions leads to lower KL-divergence growth during fine-tuning and fewer lost capabilities compared to standard abliteration. This implements the projected abliteration technique described by GrimJIM on Hugging Face, which is designed precisely to keep the good subspace intact while removing harmful subspace components.

Optimization Stability

The refined direction reduces noisy gradients during Heretic's Optuna hyperparameter search. By eliminating the unstable parallel components that oscillate between good and bad directions, orthogonalize_direction enables faster convergence and more reproducible optimization results across different model runs.

How to Enable orthogonalize_direction in Heretic

Via Configuration File

Create or modify your config.toml to include:


# config.toml

[heretic]
orthogonalize_direction = true

Heretic automatically reads this flag when instantiating the Settings object from src/heretic/config.py.

Programmatically in Python

You can also enable the flag directly in your Python code:

from heretic.config import Settings

# Load defaults and override the flag

settings = Settings()
settings.orthogonalize_direction = True

# Proceed with the main pipeline

Verifying the Effect

To confirm that orthogonalization is working, inspect the norms of your refusal directions:


# After computing refusal directions:

print("Norm before orthogonalization:", refusal_directions.norm(dim=1).mean())

if settings.orthogonalize_direction:
    # The norm reflects only the orthogonal component after processing

    print("Norm after orthogonalization:", refusal_directions.norm(dim=1).mean())

When enabled, the post-orthogonalization norm represents the magnitude of the purely orthogonal component, confirming that the parallel part has been successfully removed.

Summary

  • orthogonalize_direction removes the parallel component of refusal directions, keeping only the part orthogonal to "good" model behavior
  • The implementation in src/heretic/main.py uses vector projection to purify refusal directions before abliteration
  • Enabling this setting preserves model capabilities while specifically targeting harmful refusals
  • Benefits include improved optimization stability, lower KL-divergence, and alignment with state-of-the-art projected abliteration research
  • Configure via config.toml or the Settings class in src/heretic/config.py

Frequently Asked Questions

What is the difference between standard abliteration and projected abliteration in Heretic?

Standard abliteration subtracts the full refusal direction from model activations, which can inadvertently remove components shared with helpful behaviors. Projected abliteration, enabled by orthogonalize_direction, first removes the component of the refusal direction that overlaps with good directions, resulting in a purified vector that targets only harmful refusals without affecting the model's helpful knowledge.

Does orthogonalize_direction affect the convergence speed of Heretic's optimization?

Yes. By removing noisy parallel components that create unstable gradients during the Optuna search process, orthogonalize_direction typically leads to faster convergence and more reproducible hyperparameter optimization results. The purified direction provides cleaner gradients that allow the optimizer to find optimal abliteration parameters more efficiently.

Can I use orthogonalize_direction with any model architecture?

The orthogonalize_direction feature works with any architecture supported by Heretic's residual stream analysis, as implemented in src/heretic/model.py, because it operates on computed direction vectors rather than model-specific structures. However, the effectiveness depends on having well-defined "good" and "bad" direction sets derived from your specific dataset and model activations.

Where is the orthogonalize_direction setting defined in the Heretic codebase?

The setting is defined in src/heretic/config.py within the Settings class configuration schema (lines 179-184), and the implementation logic resides in src/heretic/main.py (lines 32-41) where the orthogonal projection is applied to refusal directions using torch.sum and F.normalize operations.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →