# Row Normalization Strategies in Heretic: A Complete Guide to LoRA Abliteration

> Explore Heretic's three row normalization strategies NONE PRE and FULL Understand how to use them to balance computational speed and weight magnitude preservation for LoRA adapters

- Repository: [Philipp Emanuel Weidmann/heretic](https://github.com/p-e-w/heretic)
- Tags: deep-dive
- Published: 2026-02-19

---

**Heretic provides three row normalization strategies—`NONE`, `PRE`, and `FULL`—that control how weight matrix rows are scaled before building LoRA-based ablation adapters, trading off between computational speed and preservation of original weight magnitudes.**

The `p-e-w/heretic` repository implements "abliteration," a technique that uses Low-Rank Adaptation (LoRA) to selectively remove capabilities from language models. Central to this process are the **row normalization strategies in Heretic**, defined in [`src/heretic/config.py`](https://github.com/p-e-w/heretic/blob/main/src/heretic/config.py) and applied in [`src/heretic/model.py`](https://github.com/p-e-w/heretic/blob/main/src/heretic/model.py), which determine whether and how the rows of a layer's weight matrix are normalized before constructing the ablation adapter.

## The Three Row Normalization Strategies

Heretic's `RowNormalization` enum offers three distinct approaches to handling weight matrix rows during abliteration. Each strategy affects how the LoRA adapter is computed and what properties of the original weights are preserved.

### NONE (No Normalization)

The `NONE` strategy uses the original weight matrix exactly as stored in the model checkpoint, performing no row normalization whatsoever.

In [`src/heretic/model.py`](https://github.com/p-e-w/heretic/blob/main/src/heretic/model.py) (lines 70-79), when `RowNormalization.NONE` is selected, the code skips all normalization steps and builds the LoRA adapter directly from the raw weight matrix `W`. This path requires the smallest LoRA rank (rank 1) and avoids any computational overhead from normalization or SVD operations.

**When to use:** Choose this default strategy for quick, directional ablations when you are not concerned about variations in row scales across the weight matrix. It provides the fastest computation path and minimal memory overhead.

### PRE (Pre-Normalization)

The `PRE` strategy normalizes weight matrix rows before computing the LoRA adapter, then re-scales the result by the original row norms to restore magnitude patterns.

During implementation in [`src/heretic/model.py`](https://github.com/p-e-w/heretic/blob/main/src/heretic/model.py), the code first applies `F.normalize(W, p=2, dim=1)` to L2-normalize each row. After computing the LoRA-B component on this normalized matrix, it multiplies by the saved row-norm vector (`lora_B = W_row_norms * lora_B`) as shown in lines 88-90. This ensures the ablation affects only the direction of weight vectors, not their magnitudes.

**When to use:** Select `PRE` when you want the ablation to be **direction-only**, removing the effect of differing row scales while preserving the original row-norm pattern. This is ideal when the relative direction of weight changes matters more than their absolute magnitude.

### FULL (Full Normalization with SVD)

The `FULL` strategy extends `PRE` by performing a low-rank SVD approximation after normalization, then renormalizing and rescaling to strictly preserve original row magnitudes.

In [`src/heretic/model.py`](https://github.com/p-e-w/heretic/blob/main/src/heretic/model.py) (lines 91-103), after normalizing the weight matrix and building the delta, the code performs an SVD-based low-rank approximation using `torch.svd_lowrank`. It then normalizes the adjusted matrix and restores the original row norms (`W = W * W_row_norms`). This path uses a higher default LoRA rank (rank 3) to maintain fidelity during the reconstruction process.

**When to use:** Use `FULL` for **norm-preserving bi-projected abliteration** when you need the overall magnitude of each weight row to remain unchanged after applying the adapter. This reduces side-effects on downstream layers and provides the most faithful representation of the intended weight change, at the cost of higher computational overhead.

## Configuring Row Normalization in Heretic

To select a strategy, instantiate the `Settings` class from [`src/heretic/config.py`](https://github.com/p-e-w/heretic/blob/main/src/heretic/config.py) with the desired `RowNormalization` enum value:

```python
from heretic.config import Settings, RowNormalization

# Strategy 1: No normalization (default)

settings_none = Settings(
    model="meta-llama/Llama-2-7b-hf",
    row_normalization=RowNormalization.NONE,
)

# Strategy 2: Pre-normalization

settings_pre = Settings(
    model="meta-llama/Llama-2-7b-hf",
    row_normalization=RowNormalization.PRE,
)

# Strategy 3: Full normalization with configurable rank

settings_full = Settings(
    model="meta-llama/Llama-2-7b-hf",
    row_normalization=RowNormalization.FULL,
    full_normalization_lora_rank=5,  # default is 3

)

```

After configuring, pass the settings to the `Heretic` class from [`src/heretic/model.py`](https://github.com/p-e-w/heretic/blob/main/src/heretic/model.py) to build the ablation adapter:

```python
from heretic.model import Heretic

heretic = Heretic(settings_full)
heretic.run_trial(["Explain quantum computing"])

```

## Implementation Details in the Source Code

The row normalization logic is centralized in [`src/heretic/model.py`](https://github.com/p-e-w/heretic/blob/main/src/heretic/model.py), with configuration definitions in [`src/heretic/config.py`](https://github.com/p-e-w/heretic/blob/main/src/heretic/config.py).

In [`src/heretic/config.py`](https://github.com/p-e-w/heretic/blob/main/src/heretic/config.py), the `RowNormalization` enum is defined at lines 22-27, while the `row_normalization` field is documented at lines 87-94. This field accepts the enum values and defaults to `RowNormalization.NONE`.

The implementation branches in [`src/heretic/model.py`](https://github.com/p-e-w/heretic/blob/main/src/heretic/model.py) handle each strategy:

- **Lines 70-79**: Compute row norms and handle the `NONE` branch, skipping normalization entirely.
- **Lines 88-90**: Implement the `PRE` strategy, applying `F.normalize(W, p=2, dim=1)` and rescaling `lora_B` by original row norms.
- **Lines 91-103**: Execute the `FULL` strategy, performing `torch.svd_lowrank` for SVD-based approximation, followed by renormalization and rescaling to preserve original magnitudes.

## Summary

- **RowNormalization.NONE** provides the fastest ablation path using raw weights and rank-1 LoRA adapters, suitable for quick directional changes without scale concerns.
- **RowNormalization.PRE** removes row scale variance by normalizing before adapter computation while preserving original norm patterns, ideal for direction-sensitive ablations.
- **RowNormalization.FULL** implements norm-preserving bi-projected abliteration using SVD-based low-rank approximation, maintaining original row magnitudes at the cost of higher computational overhead and default rank-3 adapters.
- Configure these strategies via the `row_normalization` field in [`src/heretic/config.py`](https://github.com/p-e-w/heretic/blob/main/src/heretic/config.py), with implementation logic residing in [`src/heretic/model.py`](https://github.com/p-e-w/heretic/blob/main/src/heretic/model.py).

## Frequently Asked Questions

### What is the default row normalization strategy in Heretic?

The default strategy is `RowNormalization.NONE`, which skips all normalization and uses the original weight matrix as-is. This requires the smallest LoRA rank (rank 1) and provides the fastest computation path for abliteration.

### How does the PRE normalization strategy differ from FULL?

The `PRE` strategy normalizes weight rows before computing the LoRA adapter and rescales the result by original row norms, affecting only the direction of weight changes. The `FULL` strategy adds an SVD-based low-rank approximation step and strict renormalization to preserve original row magnitudes exactly, making it norm-preserving rather than just direction-aware.

### Can I adjust the LoRA rank for the FULL normalization strategy?

Yes. When using `RowNormalization.FULL`, you can configure the `full_normalization_lora_rank` parameter in the `Settings` class (default is 3). Higher ranks provide better approximations during the SVD step but increase computational cost.

### Do I need to retrain the base model when switching normalization strategies?

No. Row normalization strategies are applied dynamically during the ablation adapter construction phase and do not permanently modify the base model weights. You can instantiate different `Settings` objects with varying `row_normalization` values and create separate `Heretic` instances to compare strategies on the same base model without any retraining.