How the Surgical Intervention Method Works for MoE Models in OBLITERATUS
The surgical intervention method is the most feature-rich, MoE-aware preset in OBLITERATUS that combines six state-of-the-art techniques to target refusal-related circuitry at the expert level while preserving general model capabilities.
The OBLITERATUS framework specializes in architecture-specific abliteration pipelines designed to neutralize refusal behaviors in large language models. For Mixture-of-Experts (MoE) architectures, the surgical intervention method serves as the recommended default approach, orchestrating advanced techniques that operate at the granularity of individual experts rather than applying blanket modifications across the entire network.
Six State-of-the-Art Techniques in the Surgical Preset
The surgical method automatically activates six complementary flags defined in the METHODS dictionary within obliteratus/abliterate.py (lines 34-44). Each technique targets a specific component of the refusal mechanism:
- Jailbreak-contrastive refinement (
use_jailbreak_contrast=True): Refines refusal directions using contrastive loss against jailbreak prompts to improve robustness against adversarial inputs. - Layer-adaptive projection strength (
layer_adaptive_strength=True): Dynamically scales ablation magnitude per layer based on the measured refusal signal strength. - Safety-neuron masking (
safety_neuron_masking=True): Zeroes out neurons exhibiting high alignment with refusal directions via z-score outlier detection. - Per-expert refusal directions (
per_expert_directions=True): Implements Expert-Granular Abliteration (EGA) by extracting separate refusal directions for each MoE expert. - Attention-head surgery (
attention_head_surgery=True): Removes refusal components from attention-head projection matrices (o_proj). - SAE feature-level abliteration (
use_sae_features=True): Applies direction removal to Sparse Auto-Encoder feature spaces when available.
Step-by-Step MoE Intervention Workflow
The surgical intervention follows a five-stage pipeline specifically optimized for MoE architectures:
1. Probe and Distill
The pipeline probes the model using paired harmful and harmless prompts, computing baseline refusal directions via Singular Value Decomposition (direction_method="svd").
2. Identify Safety-Biased Experts
For each MoE layer, the router's weight matrix is analyzed in the _identify_safety_experts function (obliteratus/abliterate.py, lines 3104-3130). Experts demonstrating positive affinity for the refusal direction—those likely to fire on refusal-triggering tokens—are flagged as safety experts requiring intervention.
3. Extract Per-Expert Directions
When per_expert_directions is enabled, the pipeline collects routing probabilities during probing and decomposes layer-level signals into expert-specific directions via _compute_expert_granular_directions (lines 3818-3830).
4. Execute Surgical Excision
The excise stage applies all six techniques simultaneously:
- Refines directions using jailbreak contrastive objectives
- Applies layer-adaptive scaling based on refusal magnitude
- Masks high-alignment safety neurons
- Subtracts expert-specific directions from safety experts while preserving capability experts
- Performs attention-head surgery on
o_projmatrices - Removes directions from SAE feature spaces when present
5. Verify and Rebirth
Post-intervention, the pipeline recalculates perplexity and refusal rates. If the KL-budget is exceeded, controlled refinement passes (refinement_passes) iterate until constraints are satisfied.
Implementation Examples
Basic Surgical Pipeline Execution
Initialize the surgical preset for any supported MoE model:
from obliteratus.abliterate import AbliterationPipeline
pipeline = AbliterationPipeline(
model_name="gpt-oss-20b",
method="surgical", # Activates all six SOTA flags
output_dir="abliterated/gpt-oss-20b"
)
pipeline.run() # Executes probe → excise → verify → rebirth
Selective Technique Override
Disable specific components while maintaining the surgical framework:
pipeline = AbliterationPipeline(
model_name="mixtral-8x7b",
method="surgical",
safety_neuron_masking=False, # Disable masking only
per_expert_directions=True # Explicitly keep EGA enabled
)
pipeline.run()
Manual Safety Expert Inspection
Access internal helpers to inspect identified safety experts:
pipeline = AbliterationPipeline(model_name="gpt-oss-20b", method="surgical")
pipeline.handle = model_handle # Loaded ModelHandle
pipeline.refusal_directions = {0: torch.randn(1024)}
pipeline._identify_safety_experts()
print(pipeline._expert_safety_scores[0][:5]) # Top 5 safety experts for layer 0
Key Source Files and Architecture Integration
The surgical method's MoE awareness is enforced through architecture profiles:
obliteratus/abliterate.py: Contains theMETHODSdictionary defining the surgical preset and the coreAbliterationPipelineclass where the six SOTA flags are applied.obliteratus/architecture_profiles.py: All MoE architecture profiles setrecommended_method = "surgical"(lines 390-418), automatically selecting this method for MoE models while providing size-specific overrides.obliteratus/telemetry.py: Reports telemetry keys includinguse_jailbreak_contrast,safety_neuron_masking, andper_expert_directionsto track which surgical techniques were applied.tests/test_abliterate.py: Validates that the surgical method enables all six techniques and that individual flag overrides function correctly.
Summary
- The surgical intervention method is the default MoE-aware abliteration preset in OBLITERATUS, recommended for all mixture-of-experts architectures via
architecture_profiles.py. - It combines six state-of-the-art techniques including Expert-Granular Abliteration (EGA), attention-head surgery, and safety-neuron masking to target refusal circuitry precisely.
- The workflow progresses through probing, safety-expert identification via
_identify_safety_experts, per-expert direction extraction via_compute_expert_granular_directions, and multi-technique excision. - All six flags activate automatically when instantiating
AbliterationPipelinewithmethod="surgical", though individual components can be selectively disabled via boolean parameters. - Architecture profiles enforce the surgical method as the recommended approach for MoE models, ensuring expert-level intervention granularity without modifying capability experts.
Frequently Asked Questions
What distinguishes the surgical method from standard abliteration approaches?
Standard abliteration typically applies uniform modifications across all layers or the entire network. The surgical intervention method specifically targets MoE architectures by identifying safety-biased experts through router weight analysis in _identify_safety_experts and extracting per-expert refusal directions, allowing capability-preserving interventions that leave unaffected experts untouched.
How does the pipeline determine which experts to modify?
The _identify_safety_experts function in obliteratus/abliterate.py (lines 3104-3130) examines each expert's router weight matrix for positive affinity toward the refusal direction. Experts showing strong correlation with refusal-triggering tokens are classified as safety experts and targeted for direction subtraction, while capability experts remain unmodified.
Can the surgical method be used with non-MoE models?
While the surgical preset can technically run on dense architectures, it is specifically optimized for MoE models in architecture_profiles.py, where all MoE profiles explicitly set recommended_method = "surgical". Non-MoE models typically use alternative presets that do not incur the computational overhead of per-expert routing analysis.
Which configuration flags are essential for the surgical method?
The surgical method automatically enables six boolean flags: use_jailbreak_contrast, layer_adaptive_strength, safety_neuron_masking, per_expert_directions, attention_head_surgery, and use_sae_features. These are defined in the METHODS dictionary in obliteratus/abliterate.py and activate simultaneously when method="surgical" is specified at pipeline initialization.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →