How Refusal Directions Are Extracted in the DISTILL Stage of OBLITERATUS

During the DISTILL stage, the AbliteratePipeline constructs a refusal subspace by applying priority-ordered mathematical extractors—including Wasserstein-optimal, LEACE, and SOM—to contrastive activation datasets, ultimately falling back to difference-of-means or standard SVD for single directions.

The OBLITERATUS framework provides a systematic approach to model editing through contrastive activation analysis. During the DISTILL stage, the pipeline identifies specific vectors in weight space that represent undesirable refusal behavior, storing these as refusal_directions and refusal_subspaces for subsequent removal by the EXCISE stage.

Stage Initialization and Model Guarding

The DISTILL stage begins with event emission and timing instrumentation. According to the OBLITERATUS source code in obliteratus/abliterate.py, the pipeline emits a "distill" event and records the start timestamp at lines 94-96:

self._emit("distill", "running", "Extracting refusal subspace...")
t0 = time.time()

For computational safety, the system implements model-size guarding. When processing small models—defined as having a hidden size less than 2048, 16 or fewer layers, or fewer than 2 billion parameters—the pipeline automatically caps the number of requested directions (n_dirs) to prevent over-ablation. This logic appears at lines 101-109 of obliteratus/abliterate.py.

Extraction Method Selection

The pipeline selects an extraction strategy based on configuration flags set during initialization. The selection hierarchy, implemented in self._distill() at lines 118-138, prioritizes methods in the following order:

  • Wasserstein-optimal extraction activates when self.use_wasserstein_optimal is enabled. This method solves a generalized eigenvalue problem that minimizes the Wasserstein-2 cost per unit of refusal removed.
  • LEACE (Linear Concept Erasure) activates when self.direction_method == "leace". This provides a closed-form optimal direction for concept erasure using generalized eigenvalue decomposition.
  • SOM (Self-Organizing Map) activates when self.direction_method == "som". This learns manifold prototypes of harmful activations and subtracts the harmless centroid.
  • Whitened SVD activates when self.use_whitened_svd is true and n_dirs > 1, provided no higher-priority extractor is selected. This performs covariance-normalized SVD to obtain orthogonal directions.
  • Difference-of-means or plain SVD serves as the default fallback for single or multiple directions respectively.

Per-Layer Direction Computation

For each layer index, the pipeline verifies that both harmful and harmless activation sets exist in self._harmful_acts and self._harmless_acts. The system then attempts extraction using the priority-ordered methods.

Wasserstein-Optimal Path

When the Wasserstein optimal extractor is active, the pipeline instantiates WassersteinOptimalExtractor and calls its extract method with the contrastive activations:

w_result = wasserstein_extractor.extract(self._harmful_acts[idx],
                                         self._harmless_acts[idx],
                                         layer_idx=idx)
self.refusal_directions[idx] = w_result.direction
self.refusal_subspaces[idx] = w_result.direction.unsqueeze(0)
norms[idx] = w_result.refusal_projection

If the configuration requests multiple directions (n_dirs > 1), the pipeline fills remaining directions using standard SVD on the difference matrix, followed by orthogonalization. This implementation appears at lines 177-190 of obliteratus/abliterate.py.

LEACE Path

For LEACE extraction, the pipeline uses LEACEExtractor to compute the refusal direction:

l_result = leace_extractor.extract(self._harmful_acts[idx],
                                   self._harmless_acts[idx],
                                   layer_idx=idx)
self.refusal_directions[idx] = l_result.direction
self.refusal_subspaces[idx] = l_result.direction.unsqueeze(0)
norms[idx] = l_result.generalized_eigenvalue

This path stores the generalized eigenvalue as the norm metric and is located at lines 215-224.

SOM Path

The SOM extractor (SOMDirectionExtractor) handles multiple directions natively by learning manifold prototypes:

som_result = som_extractor.extract(self._harmful_acts[idx],
                                   self._harmless_acts[idx],
                                   n_directions=n_dirs,
                                   layer_idx=idx)
self.refusal_subspaces[idx] = som_result.directions
self.refusal_directions[idx] = som_result.directions[0]
norms[idx] = som_result.direction_scores.sum().item() * max(som_result.coverage_score, 1e-6)

The primary direction is stored in refusal_directions, while the full subspace matrix resides in refusal_subspaces. See lines 239-252 for the implementation.

Whitened SVD and Fallback Methods

When no specialized extractor is active and n_dirs > 1, the pipeline invokes WhitenedSVDExtractor to build a whitened covariance matrix before decomposition (lines 262-270). For single-direction extraction (n_dirs == 1) without active extractors, the system falls back to the normalized difference between harmful and harmless mean activations at lines 298-301.

Post-Extraction Processing

After processing all layers, the pipeline records projection norms, emits diagnostic logs, and signals completion via a "distill" "done" event with elapsed timing information (lines 311-315). The extracted refusal_directions dictionary maps layer indices to direction vectors, while refusal_subspaces contains the full subspace matrices for the subsequent EXCISE stage.

Practical Usage Examples

To run the DISTILL stage manually with a specific extractor:

from obliteratus.abliterate import AbliteratePipeline

pipeline = AbliteratePipeline(model, tokenizer)

# Configure Wasserstein-optimal extraction

pipeline.use_wasserstein_optimal = True

# Execute extraction

pipeline._distill()
print(pipeline.refusal_directions[5])  # Access layer 5 direction

Switching to LEACE extraction requires clearing other flags:

pipeline.use_wasserstein_optimal = False
pipeline.direction_method = "leace"
pipeline._distill()

Accessing the full subspace matrix for multi-direction analysis:

layer_idx = 5
subspace = pipeline.refusal_subspaces[layer_idx]  # Shape: (n_dirs, dim)

Summary

  • The DISTILL stage analyzes contrastive activation pairs to identify refusal vectors in weight space.
  • Model-size guarding prevents over-ablation on small architectures by capping direction counts when hidden size is under 2048, layers are 16 or fewer, or parameters are under 2B.
  • Five extraction methods are available: Wasserstein-optimal, LEACE, SOM, Whitened SVD, and difference-of-means, selected via priority-ordered logic in self._distill().
  • Each layer's directions are stored in refusal_directions (primary vector) and refusal_subspaces (full matrix) for use during the EXCISE stage.
  • Key implementation resides in obliteratus/abliterate.py with extractor classes defined in the obliteratus/analysis/ directory.

Frequently Asked Questions

What is the difference between refusal_directions and refusal_subspaces?

refusal_directions stores a dictionary mapping layer indices to single direction vectors representing the primary refusal axis, while refusal_subspaces contains the full matrix of directions (shape n_dirs × dim) for layers where multiple directions were extracted. The EXCISE stage uses these subspaces to zero-out the identified behavioral components from model weights.

When should I use Wasserstein-optimal extraction versus LEACE?

Use Wasserstein-optimal extraction when you need to minimize the cost of removing refusal behavior while preserving model capabilities, as it solves for the direction that minimizes Wasserstein-2 distance per unit of refusal removed. Use LEACE when you require a closed-form optimal solution for linear concept erasure with guaranteed information removal bounds, particularly effective for single-direction extraction in classification-style tasks.

How does the pipeline prevent removing too many directions from small models?

The pipeline implements automatic guarding at lines 101-109 of obliteratus/abliterate.py. If the model has fewer than 2048 hidden dimensions, 16 or fewer layers, or less than 2 billion parameters, the system reduces the requested n_dirs to prevent over-ablation that could damage model functionality.

Where are the extractor classes implemented?

The specialized extractor implementations reside in separate modules within the analysis package:

The selection and orchestration logic remains in obliteratus/abliterate.py.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →