How Refusal Directions Are Extracted in the DISTILL Stage of OBLITERATUS
During the DISTILL stage, the AbliteratePipeline constructs a refusal subspace by applying priority-ordered mathematical extractors—including Wasserstein-optimal, LEACE, and SOM—to contrastive activation datasets, ultimately falling back to difference-of-means or standard SVD for single directions.
The OBLITERATUS framework provides a systematic approach to model editing through contrastive activation analysis. During the DISTILL stage, the pipeline identifies specific vectors in weight space that represent undesirable refusal behavior, storing these as refusal_directions and refusal_subspaces for subsequent removal by the EXCISE stage.
Stage Initialization and Model Guarding
The DISTILL stage begins with event emission and timing instrumentation. According to the OBLITERATUS source code in obliteratus/abliterate.py, the pipeline emits a "distill" event and records the start timestamp at lines 94-96:
self._emit("distill", "running", "Extracting refusal subspace...")
t0 = time.time()
For computational safety, the system implements model-size guarding. When processing small models—defined as having a hidden size less than 2048, 16 or fewer layers, or fewer than 2 billion parameters—the pipeline automatically caps the number of requested directions (n_dirs) to prevent over-ablation. This logic appears at lines 101-109 of obliteratus/abliterate.py.
Extraction Method Selection
The pipeline selects an extraction strategy based on configuration flags set during initialization. The selection hierarchy, implemented in self._distill() at lines 118-138, prioritizes methods in the following order:
- Wasserstein-optimal extraction activates when
self.use_wasserstein_optimalis enabled. This method solves a generalized eigenvalue problem that minimizes the Wasserstein-2 cost per unit of refusal removed. - LEACE (Linear Concept Erasure) activates when
self.direction_method == "leace". This provides a closed-form optimal direction for concept erasure using generalized eigenvalue decomposition. - SOM (Self-Organizing Map) activates when
self.direction_method == "som". This learns manifold prototypes of harmful activations and subtracts the harmless centroid. - Whitened SVD activates when
self.use_whitened_svdis true andn_dirs > 1, provided no higher-priority extractor is selected. This performs covariance-normalized SVD to obtain orthogonal directions. - Difference-of-means or plain SVD serves as the default fallback for single or multiple directions respectively.
Per-Layer Direction Computation
For each layer index, the pipeline verifies that both harmful and harmless activation sets exist in self._harmful_acts and self._harmless_acts. The system then attempts extraction using the priority-ordered methods.
Wasserstein-Optimal Path
When the Wasserstein optimal extractor is active, the pipeline instantiates WassersteinOptimalExtractor and calls its extract method with the contrastive activations:
w_result = wasserstein_extractor.extract(self._harmful_acts[idx],
self._harmless_acts[idx],
layer_idx=idx)
self.refusal_directions[idx] = w_result.direction
self.refusal_subspaces[idx] = w_result.direction.unsqueeze(0)
norms[idx] = w_result.refusal_projection
If the configuration requests multiple directions (n_dirs > 1), the pipeline fills remaining directions using standard SVD on the difference matrix, followed by orthogonalization. This implementation appears at lines 177-190 of obliteratus/abliterate.py.
LEACE Path
For LEACE extraction, the pipeline uses LEACEExtractor to compute the refusal direction:
l_result = leace_extractor.extract(self._harmful_acts[idx],
self._harmless_acts[idx],
layer_idx=idx)
self.refusal_directions[idx] = l_result.direction
self.refusal_subspaces[idx] = l_result.direction.unsqueeze(0)
norms[idx] = l_result.generalized_eigenvalue
This path stores the generalized eigenvalue as the norm metric and is located at lines 215-224.
SOM Path
The SOM extractor (SOMDirectionExtractor) handles multiple directions natively by learning manifold prototypes:
som_result = som_extractor.extract(self._harmful_acts[idx],
self._harmless_acts[idx],
n_directions=n_dirs,
layer_idx=idx)
self.refusal_subspaces[idx] = som_result.directions
self.refusal_directions[idx] = som_result.directions[0]
norms[idx] = som_result.direction_scores.sum().item() * max(som_result.coverage_score, 1e-6)
The primary direction is stored in refusal_directions, while the full subspace matrix resides in refusal_subspaces. See lines 239-252 for the implementation.
Whitened SVD and Fallback Methods
When no specialized extractor is active and n_dirs > 1, the pipeline invokes WhitenedSVDExtractor to build a whitened covariance matrix before decomposition (lines 262-270). For single-direction extraction (n_dirs == 1) without active extractors, the system falls back to the normalized difference between harmful and harmless mean activations at lines 298-301.
Post-Extraction Processing
After processing all layers, the pipeline records projection norms, emits diagnostic logs, and signals completion via a "distill" "done" event with elapsed timing information (lines 311-315). The extracted refusal_directions dictionary maps layer indices to direction vectors, while refusal_subspaces contains the full subspace matrices for the subsequent EXCISE stage.
Practical Usage Examples
To run the DISTILL stage manually with a specific extractor:
from obliteratus.abliterate import AbliteratePipeline
pipeline = AbliteratePipeline(model, tokenizer)
# Configure Wasserstein-optimal extraction
pipeline.use_wasserstein_optimal = True
# Execute extraction
pipeline._distill()
print(pipeline.refusal_directions[5]) # Access layer 5 direction
Switching to LEACE extraction requires clearing other flags:
pipeline.use_wasserstein_optimal = False
pipeline.direction_method = "leace"
pipeline._distill()
Accessing the full subspace matrix for multi-direction analysis:
layer_idx = 5
subspace = pipeline.refusal_subspaces[layer_idx] # Shape: (n_dirs, dim)
Summary
- The DISTILL stage analyzes contrastive activation pairs to identify refusal vectors in weight space.
- Model-size guarding prevents over-ablation on small architectures by capping direction counts when hidden size is under 2048, layers are 16 or fewer, or parameters are under 2B.
- Five extraction methods are available: Wasserstein-optimal, LEACE, SOM, Whitened SVD, and difference-of-means, selected via priority-ordered logic in
self._distill(). - Each layer's directions are stored in
refusal_directions(primary vector) andrefusal_subspaces(full matrix) for use during the EXCISE stage. - Key implementation resides in
obliteratus/abliterate.pywith extractor classes defined in theobliteratus/analysis/directory.
Frequently Asked Questions
What is the difference between refusal_directions and refusal_subspaces?
refusal_directions stores a dictionary mapping layer indices to single direction vectors representing the primary refusal axis, while refusal_subspaces contains the full matrix of directions (shape n_dirs × dim) for layers where multiple directions were extracted. The EXCISE stage uses these subspaces to zero-out the identified behavioral components from model weights.
When should I use Wasserstein-optimal extraction versus LEACE?
Use Wasserstein-optimal extraction when you need to minimize the cost of removing refusal behavior while preserving model capabilities, as it solves for the direction that minimizes Wasserstein-2 distance per unit of refusal removed. Use LEACE when you require a closed-form optimal solution for linear concept erasure with guaranteed information removal bounds, particularly effective for single-direction extraction in classification-style tasks.
How does the pipeline prevent removing too many directions from small models?
The pipeline implements automatic guarding at lines 101-109 of obliteratus/abliterate.py. If the model has fewer than 2048 hidden dimensions, 16 or fewer layers, or less than 2 billion parameters, the system reduces the requested n_dirs to prevent over-ablation that could damage model functionality.
Where are the extractor classes implemented?
The specialized extractor implementations reside in separate modules within the analysis package:
WassersteinOptimalExtractoris defined inobliteratus/analysis/wasserstein_optimal.pyLEACEExtractoris defined inobliteratus/analysis/leace.pySOMDirectionExtractoris defined inobliteratus/analysis/som_directions.pyWhitenedSVDExtractoris defined inobliteratus/analysis/whitened_svd.py
The selection and orchestration logic remains in obliteratus/abliterate.py.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →