How the Alignment Imprint Detector Informs Abliteration Decisions in OBLITERATUS
The Alignment Imprint Detector analyzes the geometric signatures of refusal directions in a model’s hidden-state space to identify its training regime (DPO, RLHF, CAI, or SFT), which then determines abliteration hyperparameters including regularization strength, refinement depth, and layer selection strategy.
The Alignment Imprint Detector is a specialized analysis module in the OBLITERATUS framework that transforms geometric properties of refusal behavior into configuration priors for surgical unalignment. By quantifying how refusal vectors are distributed across hidden layers, the detector eliminates manual hyperparameter tuning and tailors the abliteration process to the specific alignment history of the target model.
What Is the Alignment Imprint Detector?
Located in obliteratus/analysis/alignment_imprint.py, the AlignmentImprintDetector class inspects the geometry of refusal directions to classify the model’s training methodology. The module extracts six critical geometric features from the hidden-state space:
- Gini coefficient – measures concentration of refusal magnitude across layers
- Effective rank – quantifies the dimensionality of the refusal subspace
- Cross-layer smoothness – detects how gradually refusal directions shift between layers
- Tail-layer bias – identifies if refusals concentrate in final layers
- Pairwise orthogonality – measures angular dispersion between direction vectors
- Spectral decay – analyzes eigenvalue distribution of refusal covariance
The detector compares these features against empirical signatures for four training regimes: DPO (Direct Preference Optimization), RLHF (Reinforcement Learning from Human Feedback), Constitutional AI (CAI), and SFT (Supervised Fine-Tuning). This comparison produces an Alignment Imprint containing the predicted alignment method, confidence scores, and per-feature values.
Integration with the Informed Abliteration Pipeline
The detector integrates into the InformedAbliterationPipeline class defined in obliteratus/informed_pipeline.py. During the ANALYZE stage, the pipeline invokes _analyze_alignment_imprint (lines 47–71), which instantiates the detector and processes the model’s refusal directions.
The detector’s output populates an AnalysisInsights object with three key attributes:
detected_alignment_method– string identifier for the training regimealignment_confidence– scalar confidence score for the predictionalignment_probabilities– distribution over possible alignment methods
These insights persist throughout the pipeline execution and drive downstream decisions in the _derive_configuration method (lines 48–78), which translates geometric findings into concrete abliteration parameters.
How Alignment Imprints Shape Abliteration Configuration
The detected alignment method acts as a data-driven prior that configures six major aspects of the abliteration process:
Direction Extraction Strategy
The concentration of refusal vectors determines how many directions to extract. For DPO and SFT models, which exhibit highly concentrated refusal geometry (high Gini coefficient), the pipeline retains a single diff-means direction. For RLHF and Constitutional AI models, which display distributed refusals across the hidden space, the system triggers multi-direction SVD extraction to capture the full polyhedral cone structure.
Regularization Strength Calibration
The _derive_configuration method sets baseline regularization values based on the detected method:
- DPO:
0.0(minimal regularization due to clean separation) - RLHF:
0.15(moderate regularization for entangled preferences) - Constitutional AI:
0.2(highest baseline for complex constitutional layers) - SFT:
0.05(light regularization for simple supervised patterns)
If the entanglement score exceeds 0.5, the pipeline adds +0.15 to any baseline to prevent capability collateral damage during refusal removal.
Refinement Pass Allocation
The pipeline estimates self-repair risk—the model’s tendency to regenerate refusal behaviors after initial abliteration. High-risk signatures common to CAI and RLHF models trigger additional refinement passes (up to 3 total), while low-risk DPO models may complete abliteration in a single pass.
Layer Selection and Surgery Mode
Refusal Sparsity Index (RSI) comparisons determine whether to apply sparse or dense surgery. Concentrated DPO and SFT methods typically yield high RSI values, enabling sparse-surgery paths that modify only critical tail layers. Distributed RLHF refusals require dense modifications across many layers.
Bayesian Optimization Budgeting
The detected method sets the KL-divergence budget for the Bayesian optimization warm-start phase:
- DPO:
kl_budget = 0.5(permissive budget due to concentrated directions) - RLHF:
kl_budget = 0.3(conservative budget for distributed entanglement) - CAI and SFT: Intermediate values scaled to their respective geometric complexity
These budgets constrain how far the model parameters can drift during optimization, preventing catastrophic forgetting while ensuring effective refusal removal.
Practical Implementation
To run the complete informed pipeline with automatic alignment detection:
from obliteratus.informed_pipeline import InformedAbliterationPipeline
pipeline = InformedAbliterationPipeline(
model_name="meta-llama/Llama-3.1-8B-Instruct",
output_dir="abliterated_informed",
harmful_prompts=["<harmful prompt>"],
harmless_prompts=["<harmless prompt>"],
)
output_path, report = pipeline.run_informed()
print("Detected alignment method:", report.insights.detected_alignment_method)
print("Regularization used:", pipeline.regularization)
print("Refinement passes:", pipeline.refinement_passes)
For standalone analysis without full abliteration:
from obliteratus.analysis.alignment_imprint import AlignmentImprintDetector
detector = AlignmentImprintDetector()
# quick_directions maps layer index to normalized refusal direction vector
imprint = detector.detect_imprint(quick_directions)
print(AlignmentImprintDetector.format_imprint(imprint))
Summary
- The Alignment Imprint Detector in
obliteratus/analysis/alignment_imprint.pyquantifies six geometric features of refusal directions to classify models into DPO, RLHF, CAI, or SFT regimes. - The Informed Abliteration Pipeline invokes the detector during the
ANALYZEstage (lines 47–71 ofinformed_pipeline.py) and stores results in anAnalysisInsightsobject. - Detected methods drive configuration decisions: DPO triggers sparse, single-direction abliteration with minimal regularization, while RLHF and CAI prompt dense, multi-direction approaches with higher refinement passes.
- Regularization baselines range from
0.0(DPO) to0.2(CAI), with dynamic adjustments based on entanglement scores. - Bayesian optimization budgets vary by method (
0.5for DPO,0.3for RLHF) to balance refusal removal against capability preservation.
Frequently Asked Questions
What geometric features does the Alignment Imprint Detector analyze?
The detector calculates the Gini coefficient (concentration), effective rank (subspace dimensionality), cross-layer smoothness (direction stability), tail-layer bias (position concentration), pairwise orthogonality (angular dispersion), and spectral decay (eigenvalue distribution) of refusal direction vectors across the model’s hidden states.
How does the detected alignment method affect regularization strength?
Each method carries a baseline regularization value: DPO uses 0.0, RLHF uses 0.15, Constitutional AI uses 0.2, and SFT uses 0.05. If the geometric analysis reveals high entanglement (score > 0.5), the pipeline adds 0.15 to these baselines to protect non-refusal capabilities during surgery.
Can the Alignment Imprint Detector be used outside the full abliteration pipeline?
Yes. The AlignmentImprintDetector class can be instantiated independently from obliteratus/analysis/alignment_imprint.py and applied to any dictionary mapping layer indices to refusal direction vectors using the detect_imprint() method, returning a structured imprint without triggering the full abliteration workflow.
Why do DPO and RLHF models require different abliteration strategies?
DPO models exhibit concentrated refusal directions with high Gini coefficients and tail-layer bias, allowing efficient single-direction, sparse-layer surgery. RLHF models show distributed, low-orthogonality refusal patterns across many layers, necessitating multi-direction SVD extraction, dense surgery, and additional refinement passes to prevent self-repair of refusal behaviors.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →