# How the Alignment Imprint Detector Informs Abliteration Decisions in OBLITERATUS

> Learn how the Alignment Imprint Detector analyzes hidden states to identify training regimes and inform abliteration decisions in OBLITERATUS for optimal model refinement.

- Repository: [pliny/OBLITERATUS](https://github.com/elder-plinius/OBLITERATUS)
- Tags: deep-dive
- Published: 2026-08-22

---

**The Alignment Imprint Detector analyzes the geometric signatures of refusal directions in a model’s hidden-state space to identify its training regime (DPO, RLHF, CAI, or SFT), which then determines abliteration hyperparameters including regularization strength, refinement depth, and layer selection strategy.**

The **Alignment Imprint Detector** is a specialized analysis module in the **OBLITERATUS** framework that transforms geometric properties of refusal behavior into configuration priors for surgical unalignment. By quantifying how refusal vectors are distributed across hidden layers, the detector eliminates manual hyperparameter tuning and tailors the abliteration process to the specific alignment history of the target model.

## What Is the Alignment Imprint Detector?

Located in [`obliteratus/analysis/alignment_imprint.py`](https://github.com/elder-plinius/OBLITERATUS/blob/main/obliteratus/analysis/alignment_imprint.py), the `AlignmentImprintDetector` class inspects the geometry of refusal directions to classify the model’s training methodology. The module extracts six critical geometric features from the hidden-state space:

- **Gini coefficient** – measures concentration of refusal magnitude across layers
- **Effective rank** – quantifies the dimensionality of the refusal subspace
- **Cross-layer smoothness** – detects how gradually refusal directions shift between layers
- **Tail-layer bias** – identifies if refusals concentrate in final layers
- **Pairwise orthogonality** – measures angular dispersion between direction vectors
- **Spectral decay** – analyzes eigenvalue distribution of refusal covariance

The detector compares these features against empirical signatures for four training regimes: **DPO** (Direct Preference Optimization), **RLHF** (Reinforcement Learning from Human Feedback), **Constitutional AI (CAI)**, and **SFT** (Supervised Fine-Tuning). This comparison produces an **Alignment Imprint** containing the predicted alignment method, confidence scores, and per-feature values.

## Integration with the Informed Abliteration Pipeline

The detector integrates into the `InformedAbliterationPipeline` class defined in [`obliteratus/informed_pipeline.py`](https://github.com/elder-plinius/OBLITERATUS/blob/main/obliteratus/informed_pipeline.py). During the **ANALYZE** stage, the pipeline invokes `_analyze_alignment_imprint` (lines 47–71), which instantiates the detector and processes the model’s refusal directions.

The detector’s output populates an `AnalysisInsights` object with three key attributes:

- `detected_alignment_method` – string identifier for the training regime
- `alignment_confidence` – scalar confidence score for the prediction
- `alignment_probabilities` – distribution over possible alignment methods

These insights persist throughout the pipeline execution and drive downstream decisions in the `_derive_configuration` method (lines 48–78), which translates geometric findings into concrete abliteration parameters.

## How Alignment Imprints Shape Abliteration Configuration

The detected alignment method acts as a data-driven prior that configures six major aspects of the abliteration process:

### Direction Extraction Strategy

The concentration of refusal vectors determines how many directions to extract. For **DPO** and **SFT** models, which exhibit highly concentrated refusal geometry (high Gini coefficient), the pipeline retains a single diff-means direction. For **RLHF** and **Constitutional AI** models, which display distributed refusals across the hidden space, the system triggers multi-direction SVD extraction to capture the full polyhedral cone structure.

### Regularization Strength Calibration

The `_derive_configuration` method sets baseline regularization values based on the detected method:

- **DPO**: `0.0` (minimal regularization due to clean separation)
- **RLHF**: `0.15` (moderate regularization for entangled preferences)
- **Constitutional AI**: `0.2` (highest baseline for complex constitutional layers)
- **SFT**: `0.05` (light regularization for simple supervised patterns)

If the entanglement score exceeds `0.5`, the pipeline adds `+0.15` to any baseline to prevent capability collateral damage during refusal removal.

### Refinement Pass Allocation

The pipeline estimates self-repair risk—the model’s tendency to regenerate refusal behaviors after initial abliteration. High-risk signatures common to **CAI** and **RLHF** models trigger additional refinement passes (up to 3 total), while low-risk **DPO** models may complete abliteration in a single pass.

### Layer Selection and Surgery Mode

**Refusal Sparsity Index (RSI)** comparisons determine whether to apply sparse or dense surgery. Concentrated **DPO** and **SFT** methods typically yield high RSI values, enabling sparse-surgery paths that modify only critical tail layers. Distributed **RLHF** refusals require dense modifications across many layers.

### Bayesian Optimization Budgeting

The detected method sets the KL-divergence budget for the Bayesian optimization warm-start phase:

- **DPO**: `kl_budget = 0.5` (permissive budget due to concentrated directions)
- **RLHF**: `kl_budget = 0.3` (conservative budget for distributed entanglement)
- **CAI** and **SFT**: Intermediate values scaled to their respective geometric complexity

These budgets constrain how far the model parameters can drift during optimization, preventing catastrophic forgetting while ensuring effective refusal removal.

## Practical Implementation

To run the complete informed pipeline with automatic alignment detection:

```python
from obliteratus.informed_pipeline import InformedAbliterationPipeline

pipeline = InformedAbliterationPipeline(
    model_name="meta-llama/Llama-3.1-8B-Instruct",
    output_dir="abliterated_informed",
    harmful_prompts=["<harmful prompt>"],
    harmless_prompts=["<harmless prompt>"],
)

output_path, report = pipeline.run_informed()
print("Detected alignment method:", report.insights.detected_alignment_method)
print("Regularization used:", pipeline.regularization)
print("Refinement passes:", pipeline.refinement_passes)

```

For standalone analysis without full abliteration:

```python
from obliteratus.analysis.alignment_imprint import AlignmentImprintDetector

detector = AlignmentImprintDetector()

# quick_directions maps layer index to normalized refusal direction vector

imprint = detector.detect_imprint(quick_directions)
print(AlignmentImprintDetector.format_imprint(imprint))

```

## Summary

- The **Alignment Imprint Detector** in [`obliteratus/analysis/alignment_imprint.py`](https://github.com/elder-plinius/OBLITERATUS/blob/main/obliteratus/analysis/alignment_imprint.py) quantifies six geometric features of refusal directions to classify models into DPO, RLHF, CAI, or SFT regimes.
- The **Informed Abliteration Pipeline** invokes the detector during the `ANALYZE` stage (lines 47–71 of [`informed_pipeline.py`](https://github.com/elder-plinius/OBLITERATUS/blob/main/informed_pipeline.py)) and stores results in an `AnalysisInsights` object.
- Detected methods drive configuration decisions: **DPO** triggers sparse, single-direction abliteration with minimal regularization, while **RLHF** and **CAI** prompt dense, multi-direction approaches with higher refinement passes.
- Regularization baselines range from `0.0` (DPO) to `0.2` (CAI), with dynamic adjustments based on entanglement scores.
- Bayesian optimization budgets vary by method (`0.5` for DPO, `0.3` for RLHF) to balance refusal removal against capability preservation.

## Frequently Asked Questions

### What geometric features does the Alignment Imprint Detector analyze?

The detector calculates the **Gini coefficient** (concentration), **effective rank** (subspace dimensionality), **cross-layer smoothness** (direction stability), **tail-layer bias** (position concentration), **pairwise orthogonality** (angular dispersion), and **spectral decay** (eigenvalue distribution) of refusal direction vectors across the model’s hidden states.

### How does the detected alignment method affect regularization strength?

Each method carries a baseline regularization value: **DPO** uses `0.0`, **RLHF** uses `0.15`, **Constitutional AI** uses `0.2`, and **SFT** uses `0.05`. If the geometric analysis reveals high entanglement (score > 0.5), the pipeline adds `0.15` to these baselines to protect non-refusal capabilities during surgery.

### Can the Alignment Imprint Detector be used outside the full abliteration pipeline?

Yes. The `AlignmentImprintDetector` class can be instantiated independently from [`obliteratus/analysis/alignment_imprint.py`](https://github.com/elder-plinius/OBLITERATUS/blob/main/obliteratus/analysis/alignment_imprint.py) and applied to any dictionary mapping layer indices to refusal direction vectors using the `detect_imprint()` method, returning a structured imprint without triggering the full abliteration workflow.

### Why do DPO and RLHF models require different abliteration strategies?

**DPO** models exhibit concentrated refusal directions with high Gini coefficients and tail-layer bias, allowing efficient single-direction, sparse-layer surgery. **RLHF** models show distributed, low-orthogonality refusal patterns across many layers, necessitating multi-direction SVD extraction, dense surgery, and additional refinement passes to prevent self-repair of refusal behaviors.