# How Refusal Directions Are Extracted in the DISTILL Stage of OBLITERATUS

> Learn how OBLITERATUS extracts refusal directions in the DISTILL stage using mathematical extractors like Wasserstein, LEACE, and SOM for optimal subspace construction.

- Repository: [pliny/OBLITERATUS](https://github.com/elder-plinius/OBLITERATUS)
- Tags: internals
- Published: 2026-08-22

---

**During the DISTILL stage, the `AbliteratePipeline` constructs a refusal subspace by applying priority-ordered mathematical extractors—including Wasserstein-optimal, LEACE, and SOM—to contrastive activation datasets, ultimately falling back to difference-of-means or standard SVD for single directions.**

The OBLITERATUS framework provides a systematic approach to model editing through contrastive activation analysis. During the DISTILL stage, the pipeline identifies specific vectors in weight space that represent undesirable refusal behavior, storing these as `refusal_directions` and `refusal_subspaces` for subsequent removal by the EXCISE stage.

## Stage Initialization and Model Guarding

The DISTILL stage begins with event emission and timing instrumentation. According to the OBLITERATUS source code in [`obliteratus/abliterate.py`](https://github.com/elder-plinius/OBLITERATUS/blob/main/obliteratus/abliterate.py), the pipeline emits a `"distill"` event and records the start timestamp at lines 94-96:

```python
self._emit("distill", "running", "Extracting refusal subspace...")
t0 = time.time()

```

For computational safety, the system implements model-size guarding. When processing small models—defined as having a hidden size less than 2048, 16 or fewer layers, or fewer than 2 billion parameters—the pipeline automatically caps the number of requested directions (`n_dirs`) to prevent over-ablation. This logic appears at lines 101-109 of [`obliteratus/abliterate.py`](https://github.com/elder-plinius/OBLITERATUS/blob/main/obliteratus/abliterate.py).

## Extraction Method Selection

The pipeline selects an extraction strategy based on configuration flags set during initialization. The selection hierarchy, implemented in `self._distill()` at lines 118-138, prioritizes methods in the following order:

- **Wasserstein-optimal extraction** activates when `self.use_wasserstein_optimal` is enabled. This method solves a generalized eigenvalue problem that minimizes the Wasserstein-2 cost per unit of refusal removed.
- **LEACE (Linear Concept Erasure)** activates when `self.direction_method == "leace"`. This provides a closed-form optimal direction for concept erasure using generalized eigenvalue decomposition.
- **SOM (Self-Organizing Map)** activates when `self.direction_method == "som"`. This learns manifold prototypes of harmful activations and subtracts the harmless centroid.
- **Whitened SVD** activates when `self.use_whitened_svd` is true and `n_dirs > 1`, provided no higher-priority extractor is selected. This performs covariance-normalized SVD to obtain orthogonal directions.
- **Difference-of-means or plain SVD** serves as the default fallback for single or multiple directions respectively.

## Per-Layer Direction Computation

For each layer index, the pipeline verifies that both harmful and harmless activation sets exist in `self._harmful_acts` and `self._harmless_acts`. The system then attempts extraction using the priority-ordered methods.

### Wasserstein-Optimal Path

When the Wasserstein optimal extractor is active, the pipeline instantiates `WassersteinOptimalExtractor` and calls its `extract` method with the contrastive activations:

```python
w_result = wasserstein_extractor.extract(self._harmful_acts[idx],
                                         self._harmless_acts[idx],
                                         layer_idx=idx)
self.refusal_directions[idx] = w_result.direction
self.refusal_subspaces[idx] = w_result.direction.unsqueeze(0)
norms[idx] = w_result.refusal_projection

```

If the configuration requests multiple directions (`n_dirs > 1`), the pipeline fills remaining directions using standard SVD on the difference matrix, followed by orthogonalization. This implementation appears at lines 177-190 of [`obliteratus/abliterate.py`](https://github.com/elder-plinius/OBLITERATUS/blob/main/obliteratus/abliterate.py).

### LEACE Path

For LEACE extraction, the pipeline uses `LEACEExtractor` to compute the refusal direction:

```python
l_result = leace_extractor.extract(self._harmful_acts[idx],
                                   self._harmless_acts[idx],
                                   layer_idx=idx)
self.refusal_directions[idx] = l_result.direction
self.refusal_subspaces[idx] = l_result.direction.unsqueeze(0)
norms[idx] = l_result.generalized_eigenvalue

```

This path stores the generalized eigenvalue as the norm metric and is located at lines 215-224.

### SOM Path

The SOM extractor (`SOMDirectionExtractor`) handles multiple directions natively by learning manifold prototypes:

```python
som_result = som_extractor.extract(self._harmful_acts[idx],
                                   self._harmless_acts[idx],
                                   n_directions=n_dirs,
                                   layer_idx=idx)
self.refusal_subspaces[idx] = som_result.directions
self.refusal_directions[idx] = som_result.directions[0]
norms[idx] = som_result.direction_scores.sum().item() * max(som_result.coverage_score, 1e-6)

```

The primary direction is stored in `refusal_directions`, while the full subspace matrix resides in `refusal_subspaces`. See lines 239-252 for the implementation.

### Whitened SVD and Fallback Methods

When no specialized extractor is active and `n_dirs > 1`, the pipeline invokes `WhitenedSVDExtractor` to build a whitened covariance matrix before decomposition (lines 262-270). For single-direction extraction (`n_dirs == 1`) without active extractors, the system falls back to the normalized difference between harmful and harmless mean activations at lines 298-301.

## Post-Extraction Processing

After processing all layers, the pipeline records projection norms, emits diagnostic logs, and signals completion via a `"distill"` `"done"` event with elapsed timing information (lines 311-315). The extracted `refusal_directions` dictionary maps layer indices to direction vectors, while `refusal_subspaces` contains the full subspace matrices for the subsequent EXCISE stage.

## Practical Usage Examples

To run the DISTILL stage manually with a specific extractor:

```python
from obliteratus.abliterate import AbliteratePipeline

pipeline = AbliteratePipeline(model, tokenizer)

# Configure Wasserstein-optimal extraction

pipeline.use_wasserstein_optimal = True

# Execute extraction

pipeline._distill()
print(pipeline.refusal_directions[5])  # Access layer 5 direction

```

Switching to LEACE extraction requires clearing other flags:

```python
pipeline.use_wasserstein_optimal = False
pipeline.direction_method = "leace"
pipeline._distill()

```

Accessing the full subspace matrix for multi-direction analysis:

```python
layer_idx = 5
subspace = pipeline.refusal_subspaces[layer_idx]  # Shape: (n_dirs, dim)

```

## Summary

- The DISTILL stage analyzes contrastive activation pairs to identify refusal vectors in weight space.
- Model-size guarding prevents over-ablation on small architectures by capping direction counts when hidden size is under 2048, layers are 16 or fewer, or parameters are under 2B.
- Five extraction methods are available: Wasserstein-optimal, LEACE, SOM, Whitened SVD, and difference-of-means, selected via priority-ordered logic in `self._distill()`.
- Each layer's directions are stored in `refusal_directions` (primary vector) and `refusal_subspaces` (full matrix) for use during the EXCISE stage.
- Key implementation resides in [`obliteratus/abliterate.py`](https://github.com/elder-plinius/OBLITERATUS/blob/main/obliteratus/abliterate.py) with extractor classes defined in the `obliteratus/analysis/` directory.

## Frequently Asked Questions

### What is the difference between `refusal_directions` and `refusal_subspaces`?

`refusal_directions` stores a dictionary mapping layer indices to single direction vectors representing the primary refusal axis, while `refusal_subspaces` contains the full matrix of directions (shape `n_dirs × dim`) for layers where multiple directions were extracted. The EXCISE stage uses these subspaces to zero-out the identified behavioral components from model weights.

### When should I use Wasserstein-optimal extraction versus LEACE?

Use **Wasserstein-optimal extraction** when you need to minimize the cost of removing refusal behavior while preserving model capabilities, as it solves for the direction that minimizes Wasserstein-2 distance per unit of refusal removed. Use **LEACE** when you require a closed-form optimal solution for linear concept erasure with guaranteed information removal bounds, particularly effective for single-direction extraction in classification-style tasks.

### How does the pipeline prevent removing too many directions from small models?

The pipeline implements automatic guarding at lines 101-109 of [`obliteratus/abliterate.py`](https://github.com/elder-plinius/OBLITERATUS/blob/main/obliteratus/abliterate.py). If the model has fewer than 2048 hidden dimensions, 16 or fewer layers, or less than 2 billion parameters, the system reduces the requested `n_dirs` to prevent over-ablation that could damage model functionality.

### Where are the extractor classes implemented?

The specialized extractor implementations reside in separate modules within the analysis package:
- `WassersteinOptimalExtractor` is defined in [`obliteratus/analysis/wasserstein_optimal.py`](https://github.com/elder-plinius/OBLITERATUS/blob/main/obliteratus/analysis/wasserstein_optimal.py)
- `LEACEExtractor` is defined in [`obliteratus/analysis/leace.py`](https://github.com/elder-plinius/OBLITERATUS/blob/main/obliteratus/analysis/leace.py)
- `SOMDirectionExtractor` is defined in [`obliteratus/analysis/som_directions.py`](https://github.com/elder-plinius/OBLITERATUS/blob/main/obliteratus/analysis/som_directions.py)
- `WhitenedSVDExtractor` is defined in [`obliteratus/analysis/whitened_svd.py`](https://github.com/elder-plinius/OBLITERATUS/blob/main/obliteratus/analysis/whitened_svd.py)

The selection and orchestration logic remains in [`obliteratus/abliterate.py`](https://github.com/elder-plinius/OBLITERATUS/blob/main/obliteratus/abliterate.py).