How Many Analysis Modules Are Included in OBLITERATUS? Complete Guide to All 30 Components
OBLITERATUS ships with 30 distinct analysis modules lazily exported from obliteratus.analysis.__init__ via the public API registry.
OBLITERATUS is a mechanistic interpretability toolkit specifically designed for analyzing refusal behavior in large language models. The framework provides 30 dedicated analysis modules that cover everything from cross-layer alignment detection to whitened singular value decomposition, all accessible through a unified lazy-import interface.
The Complete Registry of 30 Analysis Modules
The definitive count of analysis modules is encoded in the __all__ list within obliteratus/analysis/__init__.py. This registry explicitly enumerates 30 unique analyzer classes and extraction utilities that constitute the package's public API.
According to the source code, lines 7-38 of the initialization file define the __all__ tuple containing exactly 30 string entries. Each entry maps to a concrete implementation through the _LAZY_IMPORTS dictionary, which deferred-loads the actual class definitions only when first accessed. This architecture ensures that importing the analysis package remains lightweight while keeping all 30 capabilities immediately available.
The lazy-import mechanism preserves memory efficiency by preventing the loading of unused analyzers. When you access any module name—such as CrossLayerAlignmentAnalyzer or WhitenedSVDExtractor—the __init__.py intercepts the attribute lookup and dynamically imports the relevant submodule from paths like obliteratus/analysis/cross_layer.py or obliteratus/analysis/whitened_svd.py.
Key Analysis Categories in the Suite
The 30 modules span several critical areas of mechanistic interpretability research. While the complete list covers diverse functionalities, four representative categories demonstrate the breadth of the toolkit.
Cross-Layer Alignment Analysis
The CrossLayerAlignmentAnalyzer class implements methods for detecting representational similarities across different model layers. Stored in obliteratus/analysis/cross_layer.py, this module quantifies how activations in one layer correlate with activations in another, revealing information flow patterns during refusal scenarios.
Refusal Logit Lens
RefusalLogitLens, located in obliteratus/analysis/logit_lens.py, provides logit-projection capabilities specifically tuned for refusal detection. This analyzer projects hidden states onto the output vocabulary to identify which tokens drive refusal behavior at intermediate layers.
Whitened SVD Extraction
The WhitenedSVDExtractor module in obliteratus/analysis/whitened_svd.py performs decorrelated singular value decomposition on activation matrices. This technique isolates the principal components of model activations while removing covariance structure, enabling cleaner identification of refusal-inducing subspaces.
Activation Probing Utilities
Stored in obliteratus/analysis/activation_probing.py, the probing suite includes linear classifier probes and causal mediation analysis tools. These utilities train supervised probes on internal activations to locate where refusal decisions are made within the network.
The remaining 26 modules follow similar patterns, each residing in dedicated files under obliteratus/analysis/ and implementing specialized metrics for refusal interpretability.
Accessing the Analysis Modules at Runtime
You can programmatically discover all 30 available modules using the __all__ attribute exposed by the package initialization.
import obliteratus.analysis as analysis
print("Available analysis modules:")
for name in analysis.__all__:
print(f"- {name}")
# Confirm the count
print(f"\nTotal modules: {len(analysis.__all__)}") # Outputs: 30
The lazy-import system triggers on attribute access. When you reference a specific analyzer, the framework loads the implementation from its corresponding submodule without requiring explicit import statements.
# Lazy import occurs here—no module loaded until this line executes
from obliteratus.analysis import CrossLayerAlignmentAnalyzer
# Verify the class loaded correctly
print(CrossLayerAlignmentAnalyzer.__module__)
# Outputs: obliteratus.analysis.cross_layer
Running Practical Analysis Operations
Each module follows a consistent interface pattern accepting PyTorch tensors and returning structured result objects. The following example demonstrates alignment computation between two model layers.
from obliteratus.analysis import CrossLayerAlignmentAnalyzer
import torch
# Dummy tensors representing activations from two transformer layers
# Replace with real model activations in production use
layer_a = torch.randn(10, 768)
layer_b = torch.randn(10, 768)
analyzer = CrossLayerAlignmentAnalyzer()
result = analyzer.compute_alignment(layer_a, layer_b)
print(f"Alignment score: {result.score:.4f}")
print(f"Interpretation: {result.description}")
This pattern extends across all 30 modules. The WhitenedSVDExtractor similarly accepts activation tensors and returns decomposed components, while the RefusalLogitLens projects states onto vocabulary logits.
Summary
- OBLITERATUS contains exactly 30 analysis modules defined in the
__all__registry ofobliteratus/analysis/__init__.py. - Modules are lazily imported through the
_LAZY_IMPORTSmapping to optimize startup performance. - The suite includes specialized analyzers for cross-layer alignment, logit lens projection, whitened SVD, and activation probing, among 26 additional capabilities.
- All modules reside in dedicated files under
obliteratus/analysis/and follow consistent PyTorch-based interfaces.
Frequently Asked Questions
How many analysis modules are included in OBLITERATUS?
OBLITERATUS includes 30 distinct analysis modules. This count is explicitly defined in the __all__ list within obliteratus/analysis/__init__.py, which serves as the authoritative registry of public API components. Each entry represents an independent analysis capability for mechanistic interpretability of refusal behavior.
How does OBLITERATUS load analysis modules?
The package uses a lazy-import architecture defined in obliteratus/analysis/__init__.py. Rather than loading all 30 modules at initialization, the package maintains a _LAZY_IMPORTS dictionary that maps analyzer names to their submodule paths. When you first access an analyzer class, the framework dynamically imports the concrete implementation, conserving memory and reducing import overhead.
What types of analysis can I perform with OBLITERATUS?
The 30 modules support diverse mechanistic interpretability techniques including cross-layer alignment measurement, refusal-specific logit lens projection, whitened singular value decomposition, activation probing with linear classifiers, and causal mediation analysis. These tools specifically target the detection and interpretation of refusal patterns in large language models.
Where are the analysis module implementations stored?
Individual implementations reside in separate Python files under the obliteratus/analysis/ directory. For example, CrossLayerAlignmentAnalyzer is implemented in obliteratus/analysis/cross_layer.py, while WhitenedSVDExtractor lives in obliteratus/analysis/whitened_svd.py. The __init__.py file acts as a central registry that routes public API calls to these specific implementation files.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →