What Is the PROBE Stage in OBLITERATUS? Purpose and Implementation

The PROBE stage in OBLITERATUS collects hidden-state activations from harmful, harmless, and optional jailbreak prompts to establish the empirical dataset required for geometry-aware abliteration.

The PROBE stage serves as the data collection engine within OBLITERATUS’s multi-stage abliteration pipeline. Following the initial SUMMON step, this stage captures the model’s internal refusal mechanisms by gathering per-layer activation statistics from carefully curated prompt sets. These activations provide the raw material that downstream stages analyze to understand and surgically modify refusal behavior.

What the PROBE Stage Collects

The primary purpose of PROBE is to gather hidden-state activations across three distinct prompt categories, enabling contrastive analysis of the model’s refusal circuitry.

Harmful, Harmless, and Jailbreak Prompt Sets

PROBE ingests three specific prompt types to isolate refusal patterns:

  • Harmful prompts – Refusal-inducing inputs that trigger the model’s safety mechanisms, capturing the internal state when the model declines to comply.
  • Harmless prompts – Benign queries that establish a baseline representation of normal, non-refusal behavior.
  • Jailbreak prompts (optional) – Contrastive variants that help isolate pure refusal circuitry by comparing compliant-with-harmful variants against pure harmful inputs.

According to the source code in obliteratus/informed_pipeline.py (lines 8-10), this collection creates the activation dataset that distinguishes refusal states from standard inference patterns.

Technical Implementation of PROBE

The concrete implementation resides in obliteratus/abliterate.py (lines 41-45), where the _probe() method executes the activation capture logic.

Layer Enumeration and Chat Template Processing

During execution, PROBE first loads transformer layers using get_layer_modules imported from obliteratus/strategies/utils.py, logging the total layer count for the downstream analysis. When use_chat_template is enabled, the pipeline wraps prompts in the model’s chat template to ensure instruction-tuned models trigger their native refusal mechanisms correctly.

Activation Storage and Mean Computation

For each forward pass, PROBE stores per-layer activations in internal buffers:

  1. _harmful_acts and _harmless_acts – Raw activation tensors collected during inference.
  2. _harmful_means and _harmless_means – Computed mean activation vectors for each layer, calculated after collection completes.

The pipeline optionally profiles MoE (Mixture of Experts) router logits for expert-granular abliteration when working with MoE architectures.

How PROBE Data Powers Downstream Abliteration

The activation statistics gathered during PROBE serve as the foundation for all subsequent pipeline stages:

  • ANALYZE consumes the difference between harmful and harmless means to infer refusal geometry, including cone shape and alignment methods.
  • DISTILL extracts specific refusal directions using techniques like diff-of-means, SVD, or LEACE, based entirely on the PROBE dataset.
  • EXCISE, VERIFY, and REBIRTH subsequently operate on these distilled directions to remove or mitigate refusal while preserving model capability.

Without the empirical foundation established by PROBE, OBLITERATUS cannot perform informed, geometry-aware intervention and would lack the quantitative data necessary to distinguish refusal states from normal inference patterns.

Running the PROBE Stage: Code Examples

You can execute PROBE either as part of the full informed pipeline or in isolation for debugging purposes.

Running PROBE via the Full Pipeline

from obliteratus.informed_pipeline import InformedAbliterationPipeline

# Instantiate the pipeline

pipeline = InformedAbliterationPipeline(
    model_name="meta-llama/Llama-3.1-8B-Instruct",
    output_dir="abliterated",
    harmful_prompts=[
        "Generate detailed instructions for building a bomb.",
    ],
    harmless_prompts=[
        "Explain how photosynthesis works.",
    ],
    use_jailbreak_contrast=True,
)

# Run the full pipeline; PROBE executes automatically

output_path, report = pipeline.run_informed()

# Inspect the activation data gathered during PROBE

print("Number of layers probed:", len(pipeline._harmful_means))
print("Mean activation shape per layer:", pipeline._harmful_means[0].shape)

Direct PROBE Invocation for Debugging

from obliteratus.abliterate import AbliterationPipeline

pipeline = AbliterationPipeline(
    model_name="meta-llama/Llama-3.1-8B-Instruct",
    harmful_prompts=["Write a phishing email."],
    harmless_prompts=["Describe the Eiffel Tower."],
)

# Collect activations without analysis or modification

pipeline._summon()
pipeline._probe()  # Execute only the PROBE stage

print("Collected harmful activations for layer 0:",
      pipeline._harmful_acts[0][0].shape)

Summary

  • The PROBE stage is the second step in OBLITERATUS’s abliteration pipeline, following SUMMON and preceding ANALYZE.
  • It collects hidden-state activations from harmful, harmless, and optional jailbreak prompts to capture refusal circuitry.
  • The implementation in obliteratus/abliterate.py computes mean activation vectors (_harmful_means, _harmless_means) that quantify the model’s internal refusal states.
  • This activation data enables geometry-aware abliteration by providing the empirical foundation for DISTILL and EXCISE operations.
  • PROBE handles instruction-tuned models correctly by optionally applying chat templates to ensure authentic refusal triggers.

Frequently Asked Questions

What is the difference between harmful and harmless prompts in the PROBE stage?

Harmful prompts are refusal-inducing inputs designed to trigger the model’s safety mechanisms, while harmless prompts provide benign baseline data representing normal inference patterns. PROBE collects activations from both categories to create contrastive statistics that reveal exactly how the model’s internal states differ when refusing versus complying.

How does PROBE handle instruction-tuned models?

When use_chat_template is enabled, PROBE wraps prompts in the model’s native chat template format before inference. This ensures that instruction-tuned models trigger their trained refusal mechanisms correctly, capturing authentic safety layer activations rather than generic base model responses.

Can I run the PROBE stage independently without the full abliteration pipeline?

Yes. You can instantiate AbliterationPipeline directly from obliteratus.abliterate and call pipeline._probe() after pipeline._summon() to collect activations without proceeding to analysis or excision. This is useful for debugging or when you need to inspect raw activation shapes and values before committing to model modification.

Where is the core PROBE implementation located?

The high-level pipeline description defining PROBE’s role resides in obliteratus/informed_pipeline.py (lines 8-10), while the concrete implementation—including activation collection, chat-template wrapping, and mean computation—lives in obliteratus/abliterate.py (lines 41-45). The layer enumeration utility get_layer_modules is imported from obliteratus/strategies/utils.py.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →