# Benefits of Using orthogonalize_direction in Heretic for Targeted Model Abliteration

> Learn how orthogonalize_direction in Heretic removes harmful refusal components while preserving model helpfulness for targeted abliteration. Improve model safety today.

- Repository: [Philipp Emanuel Weidmann/heretic](https://github.com/p-e-w/heretic)
- Tags: deep-dive
- Published: 2026-02-19

---

**Enabling `orthogonalize_direction` in Heretic removes the component of refusal directions that overlaps with "good" model behavior, preserving helpful capabilities while eliminating harmful refusals through projected abliteration.**

Heretic is an open-source tool for model alignment that computes *refusal directions*—vectors pointing from a model's acceptable activation space toward refusal-generating regions. When you enable the `orthogonalize_direction` setting, the algorithm applies **projected abliteration** by stripping away the parallel component of refusal directions, keeping only the orthogonal part. This precision targeting makes safety tuning less destructive and more efficient according to the p-e-w/heretic source code.

## What Does orthogonalize_direction Do in Heretic?

In the standard abliteration workflow, Heretic identifies directions in activation space that trigger refusals. However, these refusal directions often contain components that overlap with the model's "good" directions—vectors representing helpful, harmless responses.

When `orthogonalize_direction` is enabled, Heretic performs a mathematical projection that removes the component of the refusal direction that lies along the good-direction vector. The result is a purified refusal direction that exclusively represents harmful refusal behavior, without encroaching on the model's useful knowledge base.

## Implementation in the Heretic Source Code

The orthogonalization logic is implemented in [`src/heretic/main.py`](https://github.com/p-e-w/heretic/blob/main/src/heretic/main.py), where the algorithm adjusts refusal directions before they are used in abliteration:

```python
if settings.orthogonalize_direction:
    # Adjust the refusal directions so that only the component that is

    # orthogonal to the good direction is subtracted during abliteration.

    good_directions = F.normalize(good_means, p=2, dim=1)
    projection_vector = torch.sum(refusal_directions * good_directions, dim=1)
    refusal_directions = (
        refusal_directions - projection_vector.unsqueeze(1) * good_directions
    )
    refusal_directions = F.normalize(refusal_directions, p=2, dim=1)

```

This code computes the projection of refusal directions onto good directions, subtracts this parallel component, and renormalizes the result. The setting itself is defined in [`src/heretic/config.py`](https://github.com/p-e-w/heretic/blob/main/src/heretic/config.py) (lines 179-184) within the `Settings` class schema with the description: *"Whether to adjust the refusal directions so that only the component that is orthogonal to the good direction is subtracted during abliteration."*

## Key Benefits of Using orthogonalize_direction

### More Targeted Abliteration

By discarding the parallel component, `orthogonalize_direction` ensures Heretic only suppresses the specific neural pathways that drive refusals. This prevents the "collateral damage" that can occur when abliteration accidentally modifies weights responsible for legitimate reasoning capabilities.

### Preservation of Model Capabilities

The orthogonal component specifically avoids interfering with the model's ability to answer harmless prompts. When the refusal direction is purified through orthogonalization, the resulting model maintains its performance on standard benchmarks and helpful queries, resulting in fewer unintended capability degradations.

### Improved Safety-Utility Trade-off

Experimental results demonstrate that using orthogonalized directions leads to lower KL-divergence growth during fine-tuning and fewer lost capabilities compared to standard abliteration. This implements the **projected abliteration** technique described by GrimJIM on Hugging Face, which is designed precisely to keep the good subspace intact while removing harmful subspace components.

### Optimization Stability

The refined direction reduces noisy gradients during Heretic's Optuna hyperparameter search. By eliminating the unstable parallel components that oscillate between good and bad directions, `orthogonalize_direction` enables faster convergence and more reproducible optimization results across different model runs.

## How to Enable orthogonalize_direction in Heretic

### Via Configuration File

Create or modify your [`config.toml`](https://github.com/p-e-w/heretic/blob/main/config.toml) to include:

```toml

# config.toml

[heretic]
orthogonalize_direction = true

```

Heretic automatically reads this flag when instantiating the `Settings` object from [`src/heretic/config.py`](https://github.com/p-e-w/heretic/blob/main/src/heretic/config.py).

### Programmatically in Python

You can also enable the flag directly in your Python code:

```python
from heretic.config import Settings

# Load defaults and override the flag

settings = Settings()
settings.orthogonalize_direction = True

# Proceed with the main pipeline

```

### Verifying the Effect

To confirm that orthogonalization is working, inspect the norms of your refusal directions:

```python

# After computing refusal directions:

print("Norm before orthogonalization:", refusal_directions.norm(dim=1).mean())

if settings.orthogonalize_direction:
    # The norm reflects only the orthogonal component after processing

    print("Norm after orthogonalization:", refusal_directions.norm(dim=1).mean())

```

When enabled, the post-orthogonalization norm represents the magnitude of the purely orthogonal component, confirming that the parallel part has been successfully removed.

## Summary

- **`orthogonalize_direction`** removes the parallel component of refusal directions, keeping only the part orthogonal to "good" model behavior
- The implementation in [`src/heretic/main.py`](https://github.com/p-e-w/heretic/blob/main/src/heretic/main.py) uses vector projection to purify refusal directions before abliteration
- Enabling this setting preserves model capabilities while specifically targeting harmful refusals
- Benefits include improved optimization stability, lower KL-divergence, and alignment with state-of-the-art projected abliteration research
- Configure via [`config.toml`](https://github.com/p-e-w/heretic/blob/main/config.toml) or the `Settings` class in [`src/heretic/config.py`](https://github.com/p-e-w/heretic/blob/main/src/heretic/config.py)

## Frequently Asked Questions

### What is the difference between standard abliteration and projected abliteration in Heretic?

Standard abliteration subtracts the full refusal direction from model activations, which can inadvertently remove components shared with helpful behaviors. Projected abliteration, enabled by `orthogonalize_direction`, first removes the component of the refusal direction that overlaps with good directions, resulting in a purified vector that targets only harmful refusals without affecting the model's helpful knowledge.

### Does orthogonalize_direction affect the convergence speed of Heretic's optimization?

Yes. By removing noisy parallel components that create unstable gradients during the Optuna search process, `orthogonalize_direction` typically leads to faster convergence and more reproducible hyperparameter optimization results. The purified direction provides cleaner gradients that allow the optimizer to find optimal abliteration parameters more efficiently.

### Can I use orthogonalize_direction with any model architecture?

The `orthogonalize_direction` feature works with any architecture supported by Heretic's residual stream analysis, as implemented in [`src/heretic/model.py`](https://github.com/p-e-w/heretic/blob/main/src/heretic/model.py), because it operates on computed direction vectors rather than model-specific structures. However, the effectiveness depends on having well-defined "good" and "bad" direction sets derived from your specific dataset and model activations.

### Where is the orthogonalize_direction setting defined in the Heretic codebase?

The setting is defined in [`src/heretic/config.py`](https://github.com/p-e-w/heretic/blob/main/src/heretic/config.py) within the `Settings` class configuration schema (lines 179-184), and the implementation logic resides in [`src/heretic/main.py`](https://github.com/p-e-w/heretic/blob/main/src/heretic/main.py) (lines 32-41) where the orthogonal projection is applied to refusal directions using `torch.sum` and `F.normalize` operations.