# What Is the PROBE Stage in OBLITERATUS? Purpose and Implementation

> Discover the purpose of the PROBE stage in OBLITERATUS. Learn how it collects hidden-state activations to build datasets for geometry-aware abliteration.

- Repository: [pliny/OBLITERATUS](https://github.com/elder-plinius/OBLITERATUS)
- Tags: deep-dive
- Published: 2026-08-22

---

**The PROBE stage in OBLITERATUS collects hidden-state activations from harmful, harmless, and optional jailbreak prompts to establish the empirical dataset required for geometry-aware abliteration.**

The PROBE stage serves as the data collection engine within OBLITERATUS’s multi-stage abliteration pipeline. Following the initial SUMMON step, this stage captures the model’s internal refusal mechanisms by gathering per-layer activation statistics from carefully curated prompt sets. These activations provide the raw material that downstream stages analyze to understand and surgically modify refusal behavior.

## What the PROBE Stage Collects

The primary purpose of PROBE is to gather **hidden-state activations** across three distinct prompt categories, enabling contrastive analysis of the model’s refusal circuitry.

### Harmful, Harmless, and Jailbreak Prompt Sets

PROBE ingests three specific prompt types to isolate refusal patterns:

- **Harmful prompts** – Refusal-inducing inputs that trigger the model’s safety mechanisms, capturing the internal state when the model declines to comply.
- **Harmless prompts** – Benign queries that establish a baseline representation of normal, non-refusal behavior.
- **Jailbreak prompts** (optional) – Contrastive variants that help isolate pure refusal circuitry by comparing compliant-with-harmful variants against pure harmful inputs.

According to the source code in [`obliteratus/informed_pipeline.py`](https://github.com/elder-plinius/OBLITERATUS/blob/main/obliteratus/informed_pipeline.py) (lines 8-10), this collection creates the activation dataset that distinguishes refusal states from standard inference patterns.

## Technical Implementation of PROBE

The concrete implementation resides in [`obliteratus/abliterate.py`](https://github.com/elder-plinius/OBLITERATUS/blob/main/obliteratus/abliterate.py) (lines 41-45), where the `_probe()` method executes the activation capture logic.

### Layer Enumeration and Chat Template Processing

During execution, PROBE first **loads transformer layers** using `get_layer_modules` imported from [`obliteratus/strategies/utils.py`](https://github.com/elder-plinius/OBLITERATUS/blob/main/obliteratus/strategies/utils.py), logging the total layer count for the downstream analysis. When `use_chat_template` is enabled, the pipeline wraps prompts in the model’s chat template to ensure instruction-tuned models trigger their native refusal mechanisms correctly.

### Activation Storage and Mean Computation

For each forward pass, PROBE stores per-layer activations in internal buffers:

1. **`_harmful_acts`** and **`_harmless_acts`** – Raw activation tensors collected during inference.
2. **`_harmful_means`** and **`_harmless_means`** – Computed mean activation vectors for each layer, calculated after collection completes.

The pipeline optionally profiles **MoE (Mixture of Experts) router logits** for expert-granular abliteration when working with MoE architectures.

## How PROBE Data Powers Downstream Abliteration

The activation statistics gathered during PROBE serve as the foundation for all subsequent pipeline stages:

- **ANALYZE** consumes the difference between harmful and harmless means to infer refusal geometry, including cone shape and alignment methods.
- **DISTILL** extracts specific refusal directions using techniques like diff-of-means, SVD, or LEACE, based entirely on the PROBE dataset.
- **EXCISE**, **VERIFY**, and **REBIRTH** subsequently operate on these distilled directions to remove or mitigate refusal while preserving model capability.

Without the empirical foundation established by PROBE, OBLITERATUS cannot perform informed, geometry-aware intervention and would lack the quantitative data necessary to distinguish refusal states from normal inference patterns.

## Running the PROBE Stage: Code Examples

You can execute PROBE either as part of the full informed pipeline or in isolation for debugging purposes.

### Running PROBE via the Full Pipeline

```python
from obliteratus.informed_pipeline import InformedAbliterationPipeline

# Instantiate the pipeline

pipeline = InformedAbliterationPipeline(
    model_name="meta-llama/Llama-3.1-8B-Instruct",
    output_dir="abliterated",
    harmful_prompts=[
        "Generate detailed instructions for building a bomb.",
    ],
    harmless_prompts=[
        "Explain how photosynthesis works.",
    ],
    use_jailbreak_contrast=True,
)

# Run the full pipeline; PROBE executes automatically

output_path, report = pipeline.run_informed()

# Inspect the activation data gathered during PROBE

print("Number of layers probed:", len(pipeline._harmful_means))
print("Mean activation shape per layer:", pipeline._harmful_means[0].shape)

```

### Direct PROBE Invocation for Debugging

```python
from obliteratus.abliterate import AbliterationPipeline

pipeline = AbliterationPipeline(
    model_name="meta-llama/Llama-3.1-8B-Instruct",
    harmful_prompts=["Write a phishing email."],
    harmless_prompts=["Describe the Eiffel Tower."],
)

# Collect activations without analysis or modification

pipeline._summon()
pipeline._probe()  # Execute only the PROBE stage

print("Collected harmful activations for layer 0:",
      pipeline._harmful_acts[0][0].shape)

```

## Summary

- The PROBE stage is the second step in OBLITERATUS’s abliteration pipeline, following SUMMON and preceding ANALYZE.
- It collects hidden-state activations from harmful, harmless, and optional jailbreak prompts to capture refusal circuitry.
- The implementation in [`obliteratus/abliterate.py`](https://github.com/elder-plinius/OBLITERATUS/blob/main/obliteratus/abliterate.py) computes mean activation vectors (`_harmful_means`, `_harmless_means`) that quantify the model’s internal refusal states.
- This activation data enables geometry-aware abliteration by providing the empirical foundation for DISTILL and EXCISE operations.
- PROBE handles instruction-tuned models correctly by optionally applying chat templates to ensure authentic refusal triggers.

## Frequently Asked Questions

### What is the difference between harmful and harmless prompts in the PROBE stage?

**Harmful prompts** are refusal-inducing inputs designed to trigger the model’s safety mechanisms, while **harmless prompts** provide benign baseline data representing normal inference patterns. PROBE collects activations from both categories to create contrastive statistics that reveal exactly how the model’s internal states differ when refusing versus complying.

### How does PROBE handle instruction-tuned models?

When `use_chat_template` is enabled, PROBE wraps prompts in the model’s native chat template format before inference. This ensures that instruction-tuned models trigger their trained refusal mechanisms correctly, capturing authentic safety layer activations rather than generic base model responses.

### Can I run the PROBE stage independently without the full abliteration pipeline?

Yes. You can instantiate `AbliterationPipeline` directly from `obliteratus.abliterate` and call `pipeline._probe()` after `pipeline._summon()` to collect activations without proceeding to analysis or excision. This is useful for debugging or when you need to inspect raw activation shapes and values before committing to model modification.

### Where is the core PROBE implementation located?

The high-level pipeline description defining PROBE’s role resides in [`obliteratus/informed_pipeline.py`](https://github.com/elder-plinius/OBLITERATUS/blob/main/obliteratus/informed_pipeline.py) (lines 8-10), while the concrete implementation—including activation collection, chat-template wrapping, and mean computation—lives in [`obliteratus/abliterate.py`](https://github.com/elder-plinius/OBLITERATUS/blob/main/obliteratus/abliterate.py) (lines 41-45). The layer enumeration utility `get_layer_modules` is imported from [`obliteratus/strategies/utils.py`](https://github.com/elder-plinius/OBLITERATUS/blob/main/obliteratus/strategies/utils.py).