How to Use Heretic Research Features: Residual Plotting and Geometry Analysis
Heretic provides optional research utilities that visualize the refusal direction learned during abliteration through geometric statistics and animated PaCMAP projections, enabled by setting print_residual_geometry and plot_residuals to true in the Settings configuration.
The p-e-w/heretic repository includes advanced research features for analyzing how language models distinguish between harmless and harmful prompts during the abliteration process. These optional utilities generate quantitative geometry tables and visual animations of residual vectors across model layers, giving researchers deep insight into internal model representations. This guide explains how to install, configure, and use Heretic's residual plotting and geometry analysis capabilities.
Installing the Research Dependencies
The residual plotting and geometry analysis features require extra Python packages not included in the base installation. Install them using the research extra:
pip install -U heretic-llm[research]
This installs geom-median, scikit-learn, rich, PaCMAP, imageio, and matplotlib. If these packages are missing, the methods in src/heretic/analyzer.py will emit a friendly ImportError message rather than crashing.
Enabling Residual Geometry and Plotting
You can activate these features through the Settings class defined in src/heretic/config.py. The two boolean flags are print_residual_geometry and plot_residuals.
Via TOML configuration:
print_residual_geometry = true # Display cosine similarities and silhouette scores
plot_residuals = true # Generate PaCMAP animation GIF
Programmatically via Python:
from heretic.config import Settings
settings = Settings(
model="meta-llama/Meta-Llama-3-8B",
print_residual_geometry=True,
plot_residuals=True,
)
Via CLI:
The command-line interface reads these settings from config.toml. Pass a custom configuration file using --config path/to/config.toml.
Performing Residual Analysis in Python
To use the research features programmatically, instantiate the Analyzer class from src/heretic/analyzer.py with residual tensors computed from good (harmless) and bad (harmful) prompts.
from heretic.config import Settings
from heretic.model import Model
from heretic.utils import load_prompts
from heretic.analyzer import Analyzer
# Initialize configuration
settings = Settings(
model="meta-llama/Meta-Llama-3-8B",
print_residual_geometry=True,
plot_residuals=True,
)
# Load model and prompts
model = Model(settings)
good_prompts = load_prompts(settings, settings.good_prompts)
bad_prompts = load_prompts(settings, settings.bad_prompts)
# Extract residuals with shape (n_prompts, n_layers+1, hidden_dim)
good_residuals = model.get_residuals_batched(good_prompts)
bad_residuals = model.get_residuals_batched(bad_prompts)
# Analyze and visualize
analyzer = Analyzer(settings, model, good_residuals, bad_residuals)
analyzer.print_residual_geometry() # Prints statistics table
analyzer.plot_residuals() # Saves animation.gif to settings.residual_plot_path
Running this script produces two outputs. The geometry table displays cosine similarities between good and bad vectors, L2-norms, and silhouette scores for each layer using the rich library. The animation visualizes the separation between residual clouds across all layers as a GIF file.
Understanding the Research Implementation
Residual Extraction
The Model.get_residuals_batched() method in src/heretic/model.py computes hidden-state deltas between the original model and a null reference during forward passes. This yields tensors of shape (n_prompts, n_layers+1, hidden_dim) representing how each prompt affects the model's internal states.
Geometry Analysis Calculations
The Analyzer.print_residual_geometry() method performs several geometric computations per layer:
- Geometric Median: Computes the center of residual clusters using
geom_median.torch.compute_geometric_median, which is more robust to outliers than arithmetic mean. - Similarity Metrics: Calculates cosine similarities
S(g,b),S(g,r), andS(b,r)usingtorch.nn.functional.cosine_similarity. - Cluster Quality: Computes silhouette scores via
sklearn.metrics.silhouette_scoreto quantify how well good and bad residuals separate.
PaCMAP Animation Generation
The Analyzer.plot_residuals() method creates the visualization through a multi-step process:
- Dimensionality Reduction: Applies PaCMAP (
pacmap.PaCMAP) to project high-dimensional residuals to 2D, reusing each layer's projection as initialization for the next to ensure smooth transitions. - Alignment: Rotates the embedding so the line connecting good and bad medians is horizontal, making layer-to-layer comparisons intuitive.
- Frame Interpolation: Generates
N_TRANSITION_FRAMES = 20linear interpolation frames between layers for fluid animation. - Assembly: Uses
imageio.v3.imwrite()to compile PNG sequences intoanimation.gifat the path specified bysettings.residual_plot_path(default:plots/).
In src/heretic/main.py lines 445-450, the CLI automatically invokes these methods when the corresponding settings flags are enabled:
if settings.print_residual_geometry:
analyzer.print_residual_geometry()
if settings.plot_residuals:
analyzer.plot_residuals()
Summary
- Install research dependencies with
pip install heretic-llm[research] - Enable features via
print_residual_geometryandplot_residualsinSettings - Use
Analyzerclass with residual tensors fromModel.get_residuals_batched() - Outputs include a statistical geometry table and a PaCMAP animation GIF
- Implementation relies on geometric median, scikit-learn, and PaCMAP libraries
Frequently Asked Questions
What Python packages are required for Heretic's research features?
The research features require geom-median, scikit-learn, rich, PaCMAP, imageio, matplotlib, and numpy. These are installed via the heretic-llm[research] extra. If missing, the analyzer methods in src/heretic/analyzer.py will raise an ImportError with installation instructions.
How do I interpret the residual geometry table output?
The table displays cosine similarities between good (harmless) and bad (harmful) residual vectors, L2-norms measuring vector magnitudes, and silhouette scores indicating cluster separation quality per layer. Higher silhouette scores indicate better separation between the refusal and compliance directions.
Can I customize the animation output path?
Yes, the output directory is controlled by the residual_plot_path field in the Settings class. The default location is plots/, but you can specify any path when constructing your Settings instance or in the TOML configuration file.
Why does Heretic use geometric median instead of mean for residual analysis?
The geometric median, computed via geom_median.torch.compute_geometric_median, provides robustness against outliers such as extremely toxic responses that might skew arithmetic means. This ensures the central tendency of residual clusters accurately represents typical prompt behavior.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →