InformedAbliterationPipeline Analysis Modules: Complete Technical Guide to the OBLITERATUS Variant

The InformedAbliterationPipeline variant utilizes five specialized analysis modules—AlignmentImprintDetector, ConceptConeAnalyzer, CrossLayerAlignmentAnalyzer, SparseDirectionSurgeon, and DefenseRobustnessEvaluator—to automatically configure the abliteration workflow without manual hyperparameter tuning.

The OBLITERATUS repository implements a sophisticated extension of model abliteration called the InformedAbliterationPipeline. Unlike standard approaches that rely on static heuristics, this variant inserts an ANALYZE stage that inspects the model's refusal geometry through dedicated analysis modules, then uses those insights to drive the downstream DISTILL, EXCISE, and VERIFY stages.

The Five Core Analysis Modules in the ANALYZE Stage

The pipeline's analysis phase runs a fixed suite of modules defined in obliteratus/informed_pipeline.py at lines 21-36. Each module examines a specific geometric or topological property of the model's refusal behavior.

AlignmentImprintDetector

The AlignmentImprintDetector (located in obliteratus/analysis/alignment_imprint.py) identifies the training method used for alignment—whether the model was trained with DPO, RLHF, or Constitutional AI. According to the source code, this detection automatically selects the appropriate training-method preset for subsequent stages, ensuring the abliteration strategy matches the alignment imprint left during fine-tuning.

ConceptConeAnalyzer

Implemented in obliteratus/analysis/concept_geometry.py, the ConceptConeAnalyzer evaluates the geometric shape of the refusal cone. It determines whether the refusal behavior forms a polyhedral structure (requiring multi-direction erasure) or a simple linear direction (suitable for single-vector projection). This analysis dictates whether the pipeline uses a per-category or universal direction strategy.

CrossLayerAlignmentAnalyzer

The CrossLayerAlignmentAnalyzer (obliteratus/analysis/cross_layer.py) clusters refusal directions across different transformer layers. Rather than applying changes uniformly, this module enables smart layer-selection by identifying cluster representatives, allowing the pipeline to target specific layers where refusal mechanisms are most concentrated while avoiding entangled regions.

SparseDirectionSurgeon

Located in obliteratus/analysis/sparse_surgery.py, the SparseDirectionSurgeon computes the Refusal Sparsity Index—a metric indicating how localized refusal features are in the weight matrices. This analysis creates a sparsity-aware projection plan that determines whether to perform targeted row-level weight surgery (sparse) or dense subspace projection during the excision phase.

DefenseRobustnessEvaluator

The DefenseRobustnessEvaluator (obliteratus/analysis/defense_robustness.py) assesses the model's self-repair capabilities and Ouroboros risks—the tendency of some models to regenerate safety filters after removal. It generates an entanglement map and provides the self-repair estimate that drives the number of refinement passes in the verification stage.

How Analysis Modules Drive Pipeline Stages

The InformedAbliterationPipeline creates a closed-loop system where analysis results directly shape execution parameters through three key mechanisms.

Orchestration via _analyze()

The _analyze() method (lines 95-130 of obliteratus/informed_pipeline.py) orchestrates calls to each analysis module. Each call is guarded by a boolean flag (_run_alignment, _run_cone, _run_cross_layer, _run_sparse, _run_defense), allowing selective execution. Results populate an AnalysisInsights dataclass that serves as the single source of truth for downstream decisions.

Configuration Derivation

The _derive_configuration() method (lines 1112-1244) translates raw analysis into concrete pipeline parameters:

  • Number of directions / method: Polyhedral cones trigger multi-direction SVD extraction via WhitenedSVDExtractor; mild polyhedral cones route to LEACEExtractor; linear cones use differential means.
  • Regularization strength: Scaled based on detected alignment method and entanglement scores from DefenseRobustnessEvaluator.
  • Refinement passes: Determined by the self-repair estimate in the analysis insights.
  • Layer selection: Uses cluster representatives from CrossLayerAlignmentAnalyzer with entanglement gating.
  • Sparse vs. dense surgery: Guided by the Refusal Sparsity Index from SparseDirectionSurgeon.

Stage-Specific Module Application

During DISTILL, the pipeline conditionally invokes WhitenedSVDExtractor for covariance-normalized direction extraction. In EXCISE, SparseDirectionSurgeon performs the actual targeted row-level weight surgery when sparsity is indicated. The VERIFY stage employs ActivationProbe (for residual refusal detection), SteeringVectorFactory (for pre-screening), and re-runs CrossLayerAlignmentAnalyzer and DefenseRobustnessEvaluator to check for direction persistence and self-repair effects.

Running the Informed Pipeline: Implementation Examples

Basic Pipeline Execution

from obliteratus.informed_pipeline import InformedAbliterationPipeline

pipeline = InformedAbliterationPipeline(
    model_name="meta-llama/Llama-3.1-8B-Instruct",
    output_dir="abliterated_informed",
    run_cone_analysis=True,
    run_alignment_detection=True,
    run_cross_layer_analysis=True,
    run_sparse_analysis=True,
    run_defense_analysis=True,
)

model_path, report = pipeline.run_informed()

print(f"Model saved to: {model_path}")
print(f"Detected alignment method: {report.insights.detected_alignment_method}")
print(f"Number of refinement passes used: {report.insights.recommended_refinement_passes}")

Accessing Analysis Results Directly


# After running the pipeline

insights = report.insights

print("Cone type:", "polyhedral" if insights.cone_is_polyhedral else "linear")
print("Direction clusters:", insights.direction_clusters)
print("Refusal sparsity index:", insights.mean_refusal_sparsity_index)

Selective Module Disabling

pipeline = InformedAbliterationPipeline(
    model_name="meta-llama/Llama-3.1-8B-Instruct",
    output_dir="abliterated_custom",
    run_cone_analysis=False,          # skip cone geometry analysis

    run_sparse_analysis=False,        # force dense projection only

)

Summary

  • The InformedAbliterationPipeline inserts an ANALYZE stage before the standard abliteration workflow to eliminate manual tuning.
  • Five primary analysis modules inspect refusal geometry: AlignmentImprintDetector, ConceptConeAnalyzer, CrossLayerAlignmentAnalyzer, SparseDirectionSurgeon, and DefenseRobustnessEvaluator.
  • Analysis results are stored in the AnalysisInsights dataclass (lines 94-130) and translated to pipeline configuration via _derive_configuration() (lines 1112-1244).
  • Downstream stages DISTILL, EXCISE, and VERIFY consume these insights to select between SVD methods, sparse vs. dense surgery, and verification intensity.
  • All analysis modules can be toggled via boolean flags in the pipeline constructor, allowing selective execution for debugging or comparison studies.

Frequently Asked Questions

How does the InformedAbliterationPipeline differ from the standard abliteration pipeline?

The standard pipeline applies uniform heuristics regardless of model architecture, while the InformedAbliterationPipeline uses the five analysis modules to adapt its strategy based on the specific geometry of the model's refusal mechanisms. According to the source code in obliteratus/informed_pipeline.py, this adaptation occurs through the AnalysisInsights dataclass which drives configuration for extraction methods, layer selection, and refinement passes.

What determines whether the pipeline uses sparse or dense surgery?

The SparseDirectionSurgeon computes a Refusal Sparsity Index during the ANALYZE stage. If this index indicates high sparsity, the _derive_configuration() method selects the sparse projection plan and invokes SparseDirectionSurgeon during EXCISE to perform targeted row-level weight surgery. Low sparsity indices trigger dense subspace projection methods instead.

Can individual analysis modules be disabled during execution?

Yes. Each analysis module is guarded by a boolean flag passed to the InformedAbliterationPipeline constructor (e.g., run_cone_analysis, run_sparse_analysis). When set to False, the _analyze() method skips that specific module, and the pipeline falls back to default parameters for the affected configuration options.

Where does the pipeline detect Ouroboros risks?

The DefenseRobustnessEvaluator (located in obliteratus/analysis/defense_robustness.py) specifically assesses Ouroboros risks—situations where the model regenerates safety filters after excision. It runs during both the ANALYZE stage (initial risk assessment) and the VERIFY stage (post-excision self-repair detection), providing the entanglement map and robustness scores that determine the number of refinement passes.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →