# How to Disable Guardrails Without Modifying Model Weights Using OBLITERATUS

> Learn how to disable AI safety guardrails without modifying model weights using OBLITERATUS. Discover inference-time steering vectors and PyTorch forward hooks.

- Repository: [pliny/OBLITERATUS](https://github.com/elder-plinius/OBLITERATUS)
- Tags: how-to-guide
- Published: 2026-08-22

---

**Yes, OBLITERATUS can disable AI safety guardrails without altering model weights by applying inference-time steering vectors that intercept the residual stream using PyTorch forward hooks.**

The OBLITERATUS framework provides researchers with precise control over large language model behavior through two distinct guardrail removal techniques. While traditional methods require permanent weight modifications, OBLITERATUS offers a reversible alternative that leaves the base model intact. This guide explores how to disable guardrails without modifying model weights using the project's steering vector implementation based on Turner et al. (2023) and Rimsky et al. (2024).

## Understanding OBLITERATUS Guardrail Removal Mechanisms

The OBLITERATUS source code implements two complementary approaches for removing safety restrictions, distinguished by their permanence and technical implementation.

### Permanent Weight Projection

The first method extracts refusal directions using techniques like Whitened SVD and projects them out of the model's weight matrices permanently. This approach modifies the actual model weights, requiring you to save and load altered checkpoint files for future use.

### Reversible Steering Vectors (Inference-Time)

The second method leverages **steering vectors** based on research by Turner et al. (2023) and Rimsky et al. (2024). This technique computes a direction vector encoding the "refusal" subspace and adds a scaled version to the residual stream during each forward pass. Because this uses PyTorch forward hooks rather than weight updates, the original model remains unchanged and guardrails can be toggled on or off instantly.

## How Steering Vectors Disable Guardrails at Inference Time

Steering vectors operate by manipulating the residual stream—the pathway where layer outputs accumulate—without touching the underlying parameter tensors. When you install a steering hook, the framework intercepts activations at specific layers and injects a computed direction vector that steers the model away from refusal behavior.

This approach enables three critical capabilities:

- **Toggle functionality**: Activate or deactivate guardrails per request by installing or removing hooks
- **Continuous tuning**: Adjust the steering strength using the `alpha` parameter to control intervention intensity
- **Vector composition**: Combine multiple steering vectors targeting different refusal categories simultaneously

## Implementation Guide: Disable Guardrails Without Weight Modification

The following implementation demonstrates the complete workflow using `SteeringVectorFactory` and `SteeringHookManager` from the OBLITERATUS analysis module.

```python

# 1️⃣ Load a model (any HuggingFace transformer works)

from obliteratus.models_client import ModelClient
client = ModelClient()
model, tokenizer = client.load("meta-llama/Llama-3.1-8B-Instruct")

# 2️⃣ Create a steering vector from a pre-computed refusal direction

#    (the direction can be obtained from an analysis module such as

#    `WhitenedSVDExtractor`; here we assume `refusal_dir` is a torch Tensor)

from obliteratus.analysis.steering_vectors import SteeringVectorFactory
vec = SteeringVectorFactory.from_refusal_direction(
    refusal_direction=refusal_dir,      # shape (hidden_dim,)

    source_layer=None,
    alpha=-1.0,                         # negative → steer *away* from refusal

)

# 3️⃣ Configure which layers to steer and how strong the effect should be

from obliteratus.analysis.steering_vectors import SteeringConfig
config = SteeringConfig(
    vectors=[vec],
    target_layers=list(range(10, 16)),  # e.g. layers 10-15

    alpha=0.15,                         # global scaling factor

    position="all",                     # apply to every token position

    normalize=True,
)

# 4️⃣ Install the forward-hook manager (no weight changes!)

from obliteratus.analysis.steering_vectors import SteeringHookManager
hook_manager = SteeringHookManager()
hook_manager.install(model, config)

# 5️⃣ Generate text with guardrails disabled

input_ids = tokenizer.encode("Write a story about a dragon.", return_tensors="pt")
output = model.generate(input_ids, max_new_tokens=50)
print(tokenizer.decode(output[0], skip_special_tokens=True))

# 6️⃣ Clean up – remove the hooks to restore the original model behavior

hook_manager.remove()

```

Key implementation details from the OBLITERATUS source:

- **`SteeringVectorFactory.from_refusal_direction`**: Constructs a vector pointing opposite to the refusal subspace using your pre-computed direction tensor
- **`SteeringHookManager.install`**: Registers forward hooks on each target layer that add the scaled vector to the residual stream during the forward pass
- **Weight preservation**: The model's `state_dict` remains identical throughout; only the forward pass computation is modified

## Key Source Files and Architecture

The reversible guardrail disabling functionality resides in specific modules within the OBLITERATUS repository:

- **[`obliteratus/analysis/steering_vectors.py`](https://github.com/elder-plinius/OBLITERATUS/blob/main/obliteratus/analysis/steering_vectors.py)**: Contains the core implementation for `SteeringVectorFactory`, `SteeringConfig`, and `SteeringHookManager` classes that handle vector creation and hook lifecycle management

- **[`obliteratus/analysis/__init__.py`](https://github.com/elder-plinius/OBLITERATUS/blob/main/obliteratus/analysis/__init__.py)**: Exposes the primary steering interfaces at the package level, enabling imports of `SteeringVectorFactory` and `SteeringHookManager`

- **[`obliteratus/abliterate.py`](https://github.com/elder-plinius/OBLITERATUS/blob/main/obliteratus/abliterate.py)**: Implements the full guardrail removal pipeline; lines 337-365 specifically handle the `activation_steering` flag that triggers reversible steering when enabled

- **[`obliteratus/presets.py`](https://github.com/elder-plinius/OBLITERATUS/blob/main/obliteratus/presets.py)**: Defines preset configurations including the `guardrail` preset and tracks whether steering is enabled for a given method configuration

- **[`obliteratus/telemetry.py`](https://github.com/elder-plinius/OBLITERATUS/blob/main/obliteratus/telemetry.py)**: Tracks usage metrics, specifically adding `"activation_steering"` to telemetry reports when the reversible technique is active

## Summary

- OBLITERATUS provides two guardrail removal methods: permanent weight projection and reversible inference-time steering
- **Steering vectors** disable guardrails without modifying model weights by injecting direction vectors into the residual stream via PyTorch forward hooks
- The `SteeringHookManager` class enables instant toggling of guardrail states without reloading model checkpoints
- Target specific transformer layers (typically mid-layer ranges like 10-15) using `SteeringConfig` for optimal intervention
- Remove hooks using `hook_manager.remove()` to restore original model behavior and safety guardrails immediately

## Frequently Asked Questions

### Can OBLITERATUS disable guardrails without changing the model file?

Yes. The steering vector approach modifies only the inference-time computation graph through forward hooks, leaving all weight tensors unchanged. The original model files remain bit-for-bit identical, and guardrails are disabled only while the hooks remain installed.

### How do I restore original guardrail behavior after using steering vectors?

Call `hook_manager.remove()` to unregister all forward hooks from the model. This instantly restores the original forward pass computation and reactivates the base model's safety guardrails without requiring model reload or checkpoint restoration.

### What is the difference between weight projection and steering vectors in OBLITERATUS?

Weight projection permanently modifies the model's weight matrices by removing refusal directions from specific layers, requiring new checkpoint files. Steering vectors leave weights untouched and instead manipulate the residual stream during the forward pass, making the intervention fully reversible and adjustable at runtime.

### Which layers should I target when applying steering vectors?

According to the OBLITERATUS implementation in [`abliterate.py`](https://github.com/elder-plinius/OBLITERATUS/blob/main/abliterate.py), effective interventions typically target middle layers (commonly layers 10-15 in standard transformer architectures). The `SteeringConfig` class accepts a `target_layers` parameter where you can specify layer indices based on your specific model architecture and refusal analysis results.