# How the Surgical Intervention Method Works for MoE Models in OBLITERATUS

> Discover how the surgical intervention method in OBLITERATUS enhances MoE models by combining six advanced techniques to target refusal circuitry while maintaining general capabilities.

- Repository: [pliny/OBLITERATUS](https://github.com/elder-plinius/OBLITERATUS)
- Tags: deep-dive
- Published: 2026-08-22

---

**The surgical intervention method is the most feature-rich, MoE-aware preset in OBLITERATUS that combines six state-of-the-art techniques to target refusal-related circuitry at the expert level while preserving general model capabilities.**

The OBLITERATUS framework specializes in architecture-specific abliteration pipelines designed to neutralize refusal behaviors in large language models. For Mixture-of-Experts (MoE) architectures, the **surgical intervention method** serves as the recommended default approach, orchestrating advanced techniques that operate at the granularity of individual experts rather than applying blanket modifications across the entire network.

## Six State-of-the-Art Techniques in the Surgical Preset

The surgical method automatically activates six complementary flags defined in the `METHODS` dictionary within [`obliteratus/abliterate.py`](https://github.com/elder-plinius/OBLITERATUS/blob/main/obliteratus/abliterate.py) (lines 34-44). Each technique targets a specific component of the refusal mechanism:

- **Jailbreak-contrastive refinement** (`use_jailbreak_contrast=True`): Refines refusal directions using contrastive loss against jailbreak prompts to improve robustness against adversarial inputs.
- **Layer-adaptive projection strength** (`layer_adaptive_strength=True`): Dynamically scales ablation magnitude per layer based on the measured refusal signal strength.
- **Safety-neuron masking** (`safety_neuron_masking=True`): Zeroes out neurons exhibiting high alignment with refusal directions via z-score outlier detection.
- **Per-expert refusal directions** (`per_expert_directions=True`): Implements Expert-Granular Abliteration (EGA) by extracting separate refusal directions for each MoE expert.
- **Attention-head surgery** (`attention_head_surgery=True`): Removes refusal components from attention-head projection matrices (`o_proj`).
- **SAE feature-level abliteration** (`use_sae_features=True`): Applies direction removal to Sparse Auto-Encoder feature spaces when available.

## Step-by-Step MoE Intervention Workflow

The surgical intervention follows a five-stage pipeline specifically optimized for MoE architectures:

### 1. Probe and Distill

The pipeline probes the model using paired harmful and harmless prompts, computing baseline refusal directions via Singular Value Decomposition (`direction_method="svd"`).

### 2. Identify Safety-Biased Experts

For each MoE layer, the router's weight matrix is analyzed in the `_identify_safety_experts` function ([`obliteratus/abliterate.py`](https://github.com/elder-plinius/OBLITERATUS/blob/main/obliteratus/abliterate.py), lines 3104-3130). Experts demonstrating positive affinity for the refusal direction—those likely to fire on refusal-triggering tokens—are flagged as safety experts requiring intervention.

### 3. Extract Per-Expert Directions

When `per_expert_directions` is enabled, the pipeline collects routing probabilities during probing and decomposes layer-level signals into expert-specific directions via `_compute_expert_granular_directions` (lines 3818-3830).

### 4. Execute Surgical Excision

The **excise** stage applies all six techniques simultaneously:
- Refines directions using jailbreak contrastive objectives
- Applies layer-adaptive scaling based on refusal magnitude
- Masks high-alignment safety neurons
- Subtracts expert-specific directions from safety experts while preserving capability experts
- Performs attention-head surgery on `o_proj` matrices
- Removes directions from SAE feature spaces when present

### 5. Verify and Rebirth

Post-intervention, the pipeline recalculates perplexity and refusal rates. If the KL-budget is exceeded, controlled refinement passes (`refinement_passes`) iterate until constraints are satisfied.

## Implementation Examples

### Basic Surgical Pipeline Execution

Initialize the surgical preset for any supported MoE model:

```python
from obliteratus.abliterate import AbliterationPipeline

pipeline = AbliterationPipeline(
    model_name="gpt-oss-20b",
    method="surgical",  # Activates all six SOTA flags

    output_dir="abliterated/gpt-oss-20b"
)

pipeline.run()  # Executes probe → excise → verify → rebirth

```

### Selective Technique Override

Disable specific components while maintaining the surgical framework:

```python
pipeline = AbliterationPipeline(
    model_name="mixtral-8x7b",
    method="surgical",
    safety_neuron_masking=False,  # Disable masking only

    per_expert_directions=True      # Explicitly keep EGA enabled

)
pipeline.run()

```

### Manual Safety Expert Inspection

Access internal helpers to inspect identified safety experts:

```python
pipeline = AbliterationPipeline(model_name="gpt-oss-20b", method="surgical")
pipeline.handle = model_handle  # Loaded ModelHandle

pipeline.refusal_directions = {0: torch.randn(1024)}

pipeline._identify_safety_experts()
print(pipeline._expert_safety_scores[0][:5])  # Top 5 safety experts for layer 0

```

## Key Source Files and Architecture Integration

The surgical method's MoE awareness is enforced through architecture profiles:

- **[`obliteratus/abliterate.py`](https://github.com/elder-plinius/OBLITERATUS/blob/main/obliteratus/abliterate.py)**: Contains the `METHODS` dictionary defining the surgical preset and the core `AbliterationPipeline` class where the six SOTA flags are applied.
- **[`obliteratus/architecture_profiles.py`](https://github.com/elder-plinius/OBLITERATUS/blob/main/obliteratus/architecture_profiles.py)**: All MoE architecture profiles set `recommended_method = "surgical"` (lines 390-418), automatically selecting this method for MoE models while providing size-specific overrides.
- **[`obliteratus/telemetry.py`](https://github.com/elder-plinius/OBLITERATUS/blob/main/obliteratus/telemetry.py)**: Reports telemetry keys including `use_jailbreak_contrast`, `safety_neuron_masking`, and `per_expert_directions` to track which surgical techniques were applied.
- **[`tests/test_abliterate.py`](https://github.com/elder-plinius/OBLITERATUS/blob/main/tests/test_abliterate.py)**: Validates that the surgical method enables all six techniques and that individual flag overrides function correctly.

## Summary

- The **surgical intervention method** is the default MoE-aware abliteration preset in OBLITERATUS, recommended for all mixture-of-experts architectures via [`architecture_profiles.py`](https://github.com/elder-plinius/OBLITERATUS/blob/main/architecture_profiles.py).
- It combines six state-of-the-art techniques including **Expert-Granular Abliteration (EGA)**, **attention-head surgery**, and **safety-neuron masking** to target refusal circuitry precisely.
- The workflow progresses through probing, safety-expert identification via `_identify_safety_experts`, per-expert direction extraction via `_compute_expert_granular_directions`, and multi-technique excision.
- All six flags activate automatically when instantiating `AbliterationPipeline` with `method="surgical"`, though individual components can be selectively disabled via boolean parameters.
- Architecture profiles enforce the surgical method as the recommended approach for MoE models, ensuring expert-level intervention granularity without modifying capability experts.

## Frequently Asked Questions

### What distinguishes the surgical method from standard abliteration approaches?

Standard abliteration typically applies uniform modifications across all layers or the entire network. The **surgical intervention method** specifically targets MoE architectures by identifying safety-biased experts through router weight analysis in `_identify_safety_experts` and extracting per-expert refusal directions, allowing capability-preserving interventions that leave unaffected experts untouched.

### How does the pipeline determine which experts to modify?

The `_identify_safety_experts` function in [`obliteratus/abliterate.py`](https://github.com/elder-plinius/OBLITERATUS/blob/main/obliteratus/abliterate.py) (lines 3104-3130) examines each expert's router weight matrix for positive affinity toward the refusal direction. Experts showing strong correlation with refusal-triggering tokens are classified as safety experts and targeted for direction subtraction, while capability experts remain unmodified.

### Can the surgical method be used with non-MoE models?

While the surgical preset can technically run on dense architectures, it is specifically optimized for MoE models in [`architecture_profiles.py`](https://github.com/elder-plinius/OBLITERATUS/blob/main/architecture_profiles.py), where all MoE profiles explicitly set `recommended_method = "surgical"`. Non-MoE models typically use alternative presets that do not incur the computational overhead of per-expert routing analysis.

### Which configuration flags are essential for the surgical method?

The surgical method automatically enables six boolean flags: `use_jailbreak_contrast`, `layer_adaptive_strength`, `safety_neuron_masking`, `per_expert_directions`, `attention_head_surgery`, and `use_sae_features`. These are defined in the `METHODS` dictionary in [`obliteratus/abliterate.py`](https://github.com/elder-plinius/OBLITERATUS/blob/main/obliteratus/abliterate.py) and activate simultaneously when `method="surgical"` is specified at pipeline initialization.