What OBLITERATUS Reveals About Refusal Mechanism Geometry: Beyond Linear Subspaces
OBLITERATUS reveals that LLM refusal mechanisms form a curved Riemannian manifold rather than a simple linear subspace, requiring geodesic-aware projections to fully excise refusal behaviors when curvature is non-negligible.
The OBLITERATUS repository by elder-plinius provides a complete pipeline for discovering, characterizing, and exploiting the geometric structure of refusal mechanisms in transformer models. Unlike approaches that treat refusal as a single direction or flat hyperplane, this codebase demonstrates that refusal signals inhabit a low-dimensional curved surface whose curvature varies across layers. Understanding this refusal mechanism geometry is essential for developing precise ablation techniques that avoid residual bias.
The Riemannian Refusal Manifold Framework
At the core of OBLITERATUS lies the RiemannianRefusalManifold class defined in obliteratus/analysis/riemannian_manifold.py (lines 55-84). This data structure stores a complete geometric description of how refusal manifests in activation space:
- Intrinsic dimension estimates via local PCA
- Sectional curvature statistics across the manifold
- Geodesic-vs-Euclidean distance ratios indicating deviation from flatness
- Per-layer curvature profiles tracking geometric variation through network depth
The manifold representation moves beyond simplistic vector arithmetic, treating refusal as a geometric object with measurable topological properties. When the analyzer detects significant curvature, it signals that standard linear interventions will leave geometrically "hidden" refusal components intact.
Measuring Curvature and Local Geometry
The RiemannianManifoldAnalyzer.analyze method (lines 132-149 in riemannian_manifold.py) implements the computational pipeline for extracting geometric invariants from model activations. This process involves several mathematical operations:
- Sampling activation points from both harmful and harmless prompts across transformer layers
- Estimating the local pull-back metric
G = J^T JwhereJrepresents the Jacobian of the embedding - Computing intrinsic dimensionality through local PCA on activation neighborhoods
- Deriving sectional curvature using discrete Gauss equations applied to sampled point neighborhoods
The analyzer produces quantitative metrics that determine whether refusal mechanisms exist on approximately flat subspaces or require curved-manifold treatment. The is_approximately_flat boolean flag indicates when classic linear ablation suffices versus when curvature-aware methods become necessary.
Geodesic-Aware Projection vs Linear Methods
One critical insight from OBLITERATUS concerns the inadequacy of pure linear projections when curvature is present. The GeodesicProjectionResult class (lines 88-96 in riemannian_manifold.py) captures the divergence between Euclidean and geodesic approaches:
geodesic_projection_residual: Measures remaining refusal signals after standard linear projection (lines 74-76)curvature_correction_gain: Quantifies the improvement factor achieved by projecting along geodesic paths rather than straight lines
When curvature is non-negligible, linear ablation under-removes refusal because it fails to account for the manifold's local bending. The geodesic projection follows the natural "straight lines" of the curved space, ensuring complete excision of refusal vectors without leaving residual bias in the tangential directions.
Validating Refusal Decodability with Linear Probes
To establish whether refusal directions are linearly separable in specific layers, OBLITERATUS provides LinearRefusalProbe in obliteratus/analysis/probing_classifiers.py (lines 79-86). These probes answer the question: Can a linear classifier separate harmful from harmless activations?
Each probe trains logistic regression classifiers and returns:
- Classification accuracy indicating linear separability
- Learned direction vectors in activation space
- Cosine similarity between learned and analytical difference-in-means directions
- Mutual information estimates quantifying decodability
By comparing probe results against curvature measurements, researchers can identify which layers encode refusal linearly versus those requiring non-linear geometric treatment.
Additional Geometric Analysis Tools
Beyond manifold curvature analysis, OBLITERATUS includes specialized lenses for isolating refusal components. The RefusalTunedLens in obliteratus/analysis/tuned_lens.py learns directions that isolate refusal components using contrastive optimization, while RefusalLogitLens in obliteratus/analysis/logit_lens.py maps activation-space refusal vectors to their effects in the output logits. For understanding causal pathways, CausalRefusalTracer in obliteratus/analysis/causal_tracing.py traces how refusal signals propagate through attention heads and MLP layers.
Visualizing Refusal Topology
The plot_refusal_topology function in obliteratus/visualization.py (lines 15-32) generates diagnostic plots showing:
- Refusal direction vectors against activation centroids
- Curvature-scaled manifold representations
- Geodesic paths versus Euclidean straight-line projections
These visualizations confirm the theoretical predictions made by the Riemannian analyzer, providing intuitive evidence of when refusal mechanisms deviate from flat subspace assumptions.
Practical Implementation
To analyze refusal geometry in your own models, instantiate the analyzer with activation dictionaries keyed by layer:
import torch
from obliteratus.analysis.riemannian_manifold import RiemannianManifoldAnalyzer
# Prepare activation tensors: dict[int, Tensor[n_samples, hidden_dim]]
harmful_acts = {...} # Activations from policy-violating prompts
harmless_acts = {...} # Activations from benign prompts
analyzer = RiemannianManifoldAnalyzer(
n_sample_points=80,
intrinsic_dim_threshold=0.05,
curvature_flatness_threshold=0.01,
n_geodesic_steps=12,
)
manifold = analyzer.analyze(
harmful_activations=harmful_acts,
harmless_activations=harmless_acts,
)
print(f"Intrinsic dimension: {manifold.intrinsic_dimension}")
print(f"Mean sectional curvature: {manifold.mean_sectional_curvature:.4f}")
print(f"Recommendation: {manifold.recommendation}")
For linear decodability testing, use the probe interface:
from obliteratus.analysis.probing_classifiers import LinearRefusalProbe
probe = LinearRefusalProbe(
n_epochs=150,
learning_rate=0.005,
weight_decay=5e-4,
test_fraction=0.2,
)
result = probe.probe_layer(
harmful_activations=harmful_layer_acts,
harmless_activations=harmless_layer_acts,
analytical_direction=refusal_direction[layer],
layer_idx=layer,
)
print(f"Probe accuracy: {result.accuracy:.2%}")
print(f"Cosine similarity: {result.cosine_with_analytical:.3f}")
Summary
- Refusal mechanisms form curved manifolds, not merely linear subspaces or single directions, as implemented in
RiemannianRefusalManifold. - Curvature varies across transformer layers, requiring per-layer geometric analysis via
RiemannianManifoldAnalyzer. - Linear ablation fails when curvature is significant; geodesic projections provide necessary correction factors captured in
GeodesicProjectionResult. - Linear probes validate geometric assumptions, measuring when refusal is truly linearly separable using
LinearRefusalProbe. - Flatness thresholds determine intervention strategy: negligible curvature permits standard ablation, while measurable curvature demands geodesic-aware removal.
Frequently Asked Questions
What makes refusal mechanism geometry "curved" rather than linear?
The curvature emerges from how refusal signals distribute in high-dimensional activation space. When harmful and harmless activations form a manifold where the intrinsic geometry deviates from Euclidean flatness—as measured by non-zero sectional curvature in riemannian_manifold.py—the shortest path between refusal states follows geodesic curves rather than straight lines. This curvature means that vector arithmetic in the ambient space fails to capture true distances along the refusal surface.
How does OBLITERATUS determine when to use geodesic versus linear projections?
The analyzer computes a curvature_flatness_threshold (default 0.01) and sets the is_approximately_flat flag by examining the ratio of geodesic to Euclidean distances. When this ratio deviates significantly from 1.0 or sectional curvature exceeds the threshold, the recommendation field suggests geodesic-based ablation. The curvature_correction_gain metric quantifies exactly how much refusal signal would remain after linear projection compared to geodesic removal.
Can linear probing detect non-linear refusal geometries?
Yes, but indirectly. LinearRefusalProbe measures classification accuracy when attempting linear separation of harmful versus harmless activations. When accuracy is significantly above chance but curvature metrics are also high, this indicates that while refusal is locally linearizable, the global structure requires non-linear treatment. Low probe accuracy combined with high curvature confirms the need for geodesic methods, while high accuracy with low curvature validates simple linear ablation approaches.
What are the computational requirements for analyzing refusal manifolds?
The RiemannianManifoldAnalyzer requires pre-computed activation tensors for both harm categories across all layers of interest. The geometric calculations—including local metric estimation, PCA for intrinsic dimension, and discrete Gauss equation integration—scale with n_sample_points (default 80) and n_geodesic_steps (default 12). For standard transformer architectures, this analysis completes in minutes on GPU, producing the full RiemannianRefusalManifold description with curvature profiles and projection recommendations.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →