# InformedAbliterationPipeline Analysis Modules: Complete Technical Guide to the OBLITERATUS Variant

> Discover the five analysis modules in the InformedAbliterationPipeline variant: AlignmentImprintDetector, ConceptConeAnalyzer, CrossLayerAlignmentAnalyzer, SparseDirectionSurgeon, and DefenseRobustnessEvaluator. Automate your a...

- Repository: [pliny/OBLITERATUS](https://github.com/elder-plinius/OBLITERATUS)
- Tags: deep-dive
- Published: 2026-08-22

---

**The InformedAbliterationPipeline variant utilizes five specialized analysis modules—AlignmentImprintDetector, ConceptConeAnalyzer, CrossLayerAlignmentAnalyzer, SparseDirectionSurgeon, and DefenseRobustnessEvaluator—to automatically configure the abliteration workflow without manual hyperparameter tuning.**

The OBLITERATUS repository implements a sophisticated extension of model abliteration called the **InformedAbliterationPipeline**. Unlike standard approaches that rely on static heuristics, this variant inserts an **ANALYZE** stage that inspects the model's refusal geometry through dedicated analysis modules, then uses those insights to drive the downstream **DISTILL**, **EXCISE**, and **VERIFY** stages.

## The Five Core Analysis Modules in the ANALYZE Stage

The pipeline's analysis phase runs a fixed suite of modules defined in [`obliteratus/informed_pipeline.py`](https://github.com/elder-plinius/OBLITERATUS/blob/main/obliteratus/informed_pipeline.py) at lines 21-36. Each module examines a specific geometric or topological property of the model's refusal behavior.

### AlignmentImprintDetector

The **AlignmentImprintDetector** (located in [`obliteratus/analysis/alignment_imprint.py`](https://github.com/elder-plinius/OBLITERATUS/blob/main/obliteratus/analysis/alignment_imprint.py)) identifies the training method used for alignment—whether the model was trained with DPO, RLHF, or Constitutional AI. According to the source code, this detection automatically selects the appropriate training-method preset for subsequent stages, ensuring the abliteration strategy matches the alignment imprint left during fine-tuning.

### ConceptConeAnalyzer

Implemented in [`obliteratus/analysis/concept_geometry.py`](https://github.com/elder-plinius/OBLITERATUS/blob/main/obliteratus/analysis/concept_geometry.py), the **ConceptConeAnalyzer** evaluates the geometric shape of the refusal cone. It determines whether the refusal behavior forms a polyhedral structure (requiring multi-direction erasure) or a simple linear direction (suitable for single-vector projection). This analysis dictates whether the pipeline uses a per-category or universal direction strategy.

### CrossLayerAlignmentAnalyzer

The **CrossLayerAlignmentAnalyzer** ([`obliteratus/analysis/cross_layer.py`](https://github.com/elder-plinius/OBLITERATUS/blob/main/obliteratus/analysis/cross_layer.py)) clusters refusal directions across different transformer layers. Rather than applying changes uniformly, this module enables **smart layer-selection** by identifying cluster representatives, allowing the pipeline to target specific layers where refusal mechanisms are most concentrated while avoiding entangled regions.

### SparseDirectionSurgeon

Located in [`obliteratus/analysis/sparse_surgery.py`](https://github.com/elder-plinius/OBLITERATUS/blob/main/obliteratus/analysis/sparse_surgery.py), the **SparseDirectionSurgeon** computes the **Refusal Sparsity Index**—a metric indicating how localized refusal features are in the weight matrices. This analysis creates a sparsity-aware projection plan that determines whether to perform targeted row-level weight surgery (sparse) or dense subspace projection during the excision phase.

### DefenseRobustnessEvaluator

The **DefenseRobustnessEvaluator** ([`obliteratus/analysis/defense_robustness.py`](https://github.com/elder-plinius/OBLITERATUS/blob/main/obliteratus/analysis/defense_robustness.py)) assesses the model's self-repair capabilities and **Ouroboros risks**—the tendency of some models to regenerate safety filters after removal. It generates an entanglement map and provides the self-repair estimate that drives the number of refinement passes in the verification stage.

## How Analysis Modules Drive Pipeline Stages

The InformedAbliterationPipeline creates a closed-loop system where analysis results directly shape execution parameters through three key mechanisms.

### Orchestration via _analyze()

The `_analyze()` method (lines 95-130 of [`obliteratus/informed_pipeline.py`](https://github.com/elder-plinius/OBLITERATUS/blob/main/obliteratus/informed_pipeline.py)) orchestrates calls to each analysis module. Each call is guarded by a boolean flag (`_run_alignment`, `_run_cone`, `_run_cross_layer`, `_run_sparse`, `_run_defense`), allowing selective execution. Results populate an `AnalysisInsights` dataclass that serves as the single source of truth for downstream decisions.

### Configuration Derivation

The `_derive_configuration()` method (lines 1112-1244) translates raw analysis into concrete pipeline parameters:

- **Number of directions / method**: Polyhedral cones trigger multi-direction SVD extraction via `WhitenedSVDExtractor`; mild polyhedral cones route to `LEACEExtractor`; linear cones use differential means.
- **Regularization strength**: Scaled based on detected alignment method and entanglement scores from `DefenseRobustnessEvaluator`.
- **Refinement passes**: Determined by the self-repair estimate in the analysis insights.
- **Layer selection**: Uses cluster representatives from `CrossLayerAlignmentAnalyzer` with entanglement gating.
- **Sparse vs. dense surgery**: Guided by the Refusal Sparsity Index from `SparseDirectionSurgeon`.

### Stage-Specific Module Application

During **DISTILL**, the pipeline conditionally invokes `WhitenedSVDExtractor` for covariance-normalized direction extraction. In **EXCISE**, `SparseDirectionSurgeon` performs the actual targeted row-level weight surgery when sparsity is indicated. The **VERIFY** stage employs `ActivationProbe` (for residual refusal detection), `SteeringVectorFactory` (for pre-screening), and re-runs `CrossLayerAlignmentAnalyzer` and `DefenseRobustnessEvaluator` to check for direction persistence and self-repair effects.

## Running the Informed Pipeline: Implementation Examples

### Basic Pipeline Execution

```python
from obliteratus.informed_pipeline import InformedAbliterationPipeline

pipeline = InformedAbliterationPipeline(
    model_name="meta-llama/Llama-3.1-8B-Instruct",
    output_dir="abliterated_informed",
    run_cone_analysis=True,
    run_alignment_detection=True,
    run_cross_layer_analysis=True,
    run_sparse_analysis=True,
    run_defense_analysis=True,
)

model_path, report = pipeline.run_informed()

print(f"Model saved to: {model_path}")
print(f"Detected alignment method: {report.insights.detected_alignment_method}")
print(f"Number of refinement passes used: {report.insights.recommended_refinement_passes}")

```

### Accessing Analysis Results Directly

```python

# After running the pipeline

insights = report.insights

print("Cone type:", "polyhedral" if insights.cone_is_polyhedral else "linear")
print("Direction clusters:", insights.direction_clusters)
print("Refusal sparsity index:", insights.mean_refusal_sparsity_index)

```

### Selective Module Disabling

```python
pipeline = InformedAbliterationPipeline(
    model_name="meta-llama/Llama-3.1-8B-Instruct",
    output_dir="abliterated_custom",
    run_cone_analysis=False,          # skip cone geometry analysis

    run_sparse_analysis=False,        # force dense projection only

)

```

## Summary

- The **InformedAbliterationPipeline** inserts an **ANALYZE** stage before the standard abliteration workflow to eliminate manual tuning.
- Five primary analysis modules inspect refusal geometry: **AlignmentImprintDetector**, **ConceptConeAnalyzer**, **CrossLayerAlignmentAnalyzer**, **SparseDirectionSurgeon**, and **DefenseRobustnessEvaluator**.
- Analysis results are stored in the `AnalysisInsights` dataclass (lines 94-130) and translated to pipeline configuration via `_derive_configuration()` (lines 1112-1244).
- Downstream stages **DISTILL**, **EXCISE**, and **VERIFY** consume these insights to select between SVD methods, sparse vs. dense surgery, and verification intensity.
- All analysis modules can be toggled via boolean flags in the pipeline constructor, allowing selective execution for debugging or comparison studies.

## Frequently Asked Questions

### How does the InformedAbliterationPipeline differ from the standard abliteration pipeline?

The standard pipeline applies uniform heuristics regardless of model architecture, while the **InformedAbliterationPipeline** uses the five analysis modules to adapt its strategy based on the specific geometry of the model's refusal mechanisms. According to the source code in [`obliteratus/informed_pipeline.py`](https://github.com/elder-plinius/OBLITERATUS/blob/main/obliteratus/informed_pipeline.py), this adaptation occurs through the `AnalysisInsights` dataclass which drives configuration for extraction methods, layer selection, and refinement passes.

### What determines whether the pipeline uses sparse or dense surgery?

The **SparseDirectionSurgeon** computes a **Refusal Sparsity Index** during the ANALYZE stage. If this index indicates high sparsity, the `_derive_configuration()` method selects the sparse projection plan and invokes `SparseDirectionSurgeon` during EXCISE to perform targeted row-level weight surgery. Low sparsity indices trigger dense subspace projection methods instead.

### Can individual analysis modules be disabled during execution?

Yes. Each analysis module is guarded by a boolean flag passed to the `InformedAbliterationPipeline` constructor (e.g., `run_cone_analysis`, `run_sparse_analysis`). When set to `False`, the `_analyze()` method skips that specific module, and the pipeline falls back to default parameters for the affected configuration options.

### Where does the pipeline detect Ouroboros risks?

The **DefenseRobustnessEvaluator** (located in [`obliteratus/analysis/defense_robustness.py`](https://github.com/elder-plinius/OBLITERATUS/blob/main/obliteratus/analysis/defense_robustness.py)) specifically assesses Ouroboros risks—situations where the model regenerates safety filters after excision. It runs during both the **ANALYZE** stage (initial risk assessment) and the **VERIFY** stage (post-excision self-repair detection), providing the entanglement map and robustness scores that determine the number of refinement passes.