# Understanding direction_index and Layer Selection in Heretic's Ablation Scope

> Master Heretic's ablation scope with precise control over direction_index and layer selection. Learn how these parameters fine-tune your model's refusal directions for optimal results.

- Repository: [Philipp Emanuel Weidmann/heretic](https://github.com/p-e-w/heretic)
- Tags: deep-dive
- Published: 2026-02-19

---

**The `direction_index` parameter selects either a globally interpolated refusal direction applied to all layers or defers to per-layer directions, controlled by the `direction_scope` setting in Heretic's abliteration process.**

Heretic is an open-source tool for "abliterating" language models—systematically nudging layers toward refusing harmful outputs while preserving helpful capabilities. The ablation process relies on pre-computed refusal directions that can be applied either uniformly across all layers or tailored individually per layer. Understanding how `direction_index` and `direction_scope` interact is essential for configuring effective ablation trials in the p-e-w/heretic repository.

## How direction_index and direction_scope Define the Refusal Strategy

Heretic's abliteration requires a **refusal direction** for every modified layer. Two trial parameters control which direction vector is used and how it maps to the model's architecture.

### Global vs. Per-Layer Scope

The `direction_scope` parameter determines whether the trial uses a single direction for all layers or individual directions per layer:

- **`"global"`**: One interpolated direction is computed and applied uniformly to every layer
- **`"per layer"`**: Each layer uses its own pre-computed direction from the refusal direction sequence

When `direction_scope` is set to `"per layer"`, the `direction_index` parameter is forced to `None`, disabling interpolation and enabling layer-specific selection.

### Trial Parameter Sampling in main.py

In [`src/heretic/main.py`](https://github.com/p-e-w/heretic/blob/main/src/heretic/main.py), the Optuna trial definition samples these values at lines 64-70 and 81-89:

```python
direction_scope = trial.suggest_categorical(
    "direction_scope",
    [
        "global",
        "per layer",
    ],
)

# ... later in the trial ...

direction_index = trial.suggest_float(
    "direction_index",
    0.4 * last_layer_index,
    0.9 * last_layer_index,
)
if direction_scope == "per layer":
    direction_index = None

```

The float sampling range `[0.4·L‑1, 0.9·L‑1]` (where `L` is the number of layers) ensures the index stays within valid bounds of the pre-computed direction sequence. These parameters are stored on the trial object and passed to `model.abliterate()` at line 540.

## Core Implementation in model.py

The `abliterate` method in [`src/heretic/model.py`](https://github.com/p-e-w/heretic/blob/main/src/heretic/model.py) implements the logic that consumes these parameters and applies directions to model components.

### Interpolating the Global Direction

When `direction_index` is not `None`, the code linearly interpolates between the two nearest refusal directions (lines 88-98):

```python
if direction_index is None:
    refusal_direction = None
else:
    weight, index = math.modf(direction_index + 1)
    refusal_direction = F.normalize(
        refusal_directions[int(index)].lerp(
            refusal_directions[int(index) + 1],
            weight,
        ),
        p=2,
        dim=0,
    )

```

The `math.modf()` call splits the float into fractional and integer components, enabling smooth interpolation between discrete direction vectors using PyTorch's `lerp()` function.

### Layer Iteration and Direction Assignment

The method iterates over every layer and applies the selected direction (lines 100-108):

```python
for layer_index in range(len(self.get_layers())):
    for component, modules in self.get_layer_modules(layer_index).items():
        # ... component handling ...

        if refusal_direction is None:
            layer_refusal_direction = refusal_directions[layer_index + 1]
        else:
            layer_refusal_direction = refusal_direction

```

In **per-layer mode** (`refusal_direction is None`), the code pulls the specific vector for that layer from `refusal_directions[layer_index + 1]`. In **global mode**, it reuses the interpolated `refusal_direction` for all layers. The `+ 1` offset accounts for the sequence indexing in the refusal directions tensor.

## Running Heretic with Different Direction Scopes

### Using a Global Direction

To apply the same interpolated refusal direction across all layers:

```bash
heretic run \
    --trials 30 \
    --direction-scope global \
    --direction-index 5.2

```

All layers will be nudged along the identical interpolated vector derived from position 5.2 in the refusal direction sequence.

### Using Per-Layer Directions

To allow each layer to use its own specific refusal direction:

```bash
heretic run \
    --trials 30 \
    --direction-scope per-layer

```

The driver automatically forces `direction_index` to `None`, and the model selects individual directions from `refusal_directions[layer_index + 1]` for each layer.

### Debugging Direction Selection

To inspect the chosen parameters during optimization:

```python
def objective(trial: Trial) -> float:
    # ... sampling logic ...

    print(f"Scope: {direction_scope}")
    print(f"Index: {direction_index}")  # None when per-layer

    # ... rest of objective ...

```

These values correspond exactly to the parameters consumed by `model.abliterate()`.

## Summary

- **`direction_scope`** chooses between `"global"` (uniform direction) and `"per layer"` (individual directions)
- **`direction_index`** selects a specific point in the refusal direction sequence via linear interpolation when using global scope
- When scope is `"per layer"`, `direction_index` is forced to `None` to disable interpolation
- In [`src/heretic/main.py`](https://github.com/p-e-w/heretic/blob/main/src/heretic/main.py) (lines 64-89), Optuna samples these parameters before passing them to the model
- In [`src/heretic/model.py`](https://github.com/p-e-w/heretic/blob/main/src/heretic/model.py) (lines 88-108), the code either interpolates a global vector or selects layer-specific vectors based on the `direction_index` value
- The layer loop in [`model.py`](https://github.com/p-e-w/heretic/blob/main/model.py) uses `layer_index + 1` to index into the pre-computed refusal directions array

## Frequently Asked Questions

### What happens when direction_index is set to None?

When `direction_index` is `None`, Heretic operates in per-layer mode. Each transformer layer receives its own refusal direction from the pre-computed sequence at index `layer_index + 1`. This allows the abliteration to tailor the refusal vector to each layer's specific representation space rather than applying a one-size-fits-all direction.

### How does the interpolation calculation work for global directions?

The code uses `math.modf(direction_index + 1)` to separate the float into integer and fractional parts. The integer component selects the base direction vector, while the fractional component determines the interpolation weight between that vector and the next one using PyTorch's `lerp()` function. The result is L2-normalized before application to ensure consistent magnitude across different index values.

### Can I manually override direction_index from the command line?

Yes, you can pass `--direction-index` as a command-line argument when running `heretic run`. However, this only takes effect when `--direction-scope global` is also specified. If you specify `--direction-scope per-layer`, the CLI ignores any manually provided index and forces the value to `None` to ensure each layer uses its specific direction vector.

### Which layers are affected when using per-layer mode?

In per-layer mode, every layer returned by `self.get_layers()` receives a refusal direction, specifically from `refusal_directions[layer_index + 1]`. The `+ 1` offset means layer 0 uses refusal direction index 1, layer 1 uses index 2, and so on. This indexing convention aligns with how the refusal directions tensor is structured in the pre-computed cache, ensuring each layer modification corresponds to the correct representation steering vector.