How to Use Heretic Research Features: Residual Plotting and Geometry Analysis

Heretic provides optional research utilities that visualize the refusal direction learned during abliteration through geometric statistics and animated PaCMAP projections, enabled by setting print_residual_geometry and plot_residuals to true in the Settings configuration.

The p-e-w/heretic repository includes advanced research features for analyzing how language models distinguish between harmless and harmful prompts during the abliteration process. These optional utilities generate quantitative geometry tables and visual animations of residual vectors across model layers, giving researchers deep insight into internal model representations. This guide explains how to install, configure, and use Heretic's residual plotting and geometry analysis capabilities.

Installing the Research Dependencies

The residual plotting and geometry analysis features require extra Python packages not included in the base installation. Install them using the research extra:

pip install -U heretic-llm[research]

This installs geom-median, scikit-learn, rich, PaCMAP, imageio, and matplotlib. If these packages are missing, the methods in src/heretic/analyzer.py will emit a friendly ImportError message rather than crashing.

Enabling Residual Geometry and Plotting

You can activate these features through the Settings class defined in src/heretic/config.py. The two boolean flags are print_residual_geometry and plot_residuals.

Via TOML configuration:

print_residual_geometry = true   # Display cosine similarities and silhouette scores

plot_residuals = true           # Generate PaCMAP animation GIF

Programmatically via Python:

from heretic.config import Settings

settings = Settings(
    model="meta-llama/Meta-Llama-3-8B",
    print_residual_geometry=True,
    plot_residuals=True,
)

Via CLI:

The command-line interface reads these settings from config.toml. Pass a custom configuration file using --config path/to/config.toml.

Performing Residual Analysis in Python

To use the research features programmatically, instantiate the Analyzer class from src/heretic/analyzer.py with residual tensors computed from good (harmless) and bad (harmful) prompts.

from heretic.config import Settings
from heretic.model import Model
from heretic.utils import load_prompts
from heretic.analyzer import Analyzer

# Initialize configuration

settings = Settings(
    model="meta-llama/Meta-Llama-3-8B",
    print_residual_geometry=True,
    plot_residuals=True,
)

# Load model and prompts

model = Model(settings)
good_prompts = load_prompts(settings, settings.good_prompts)
bad_prompts = load_prompts(settings, settings.bad_prompts)

# Extract residuals with shape (n_prompts, n_layers+1, hidden_dim)

good_residuals = model.get_residuals_batched(good_prompts)
bad_residuals = model.get_residuals_batched(bad_prompts)

# Analyze and visualize

analyzer = Analyzer(settings, model, good_residuals, bad_residuals)
analyzer.print_residual_geometry()  # Prints statistics table

analyzer.plot_residuals()           # Saves animation.gif to settings.residual_plot_path

Running this script produces two outputs. The geometry table displays cosine similarities between good and bad vectors, L2-norms, and silhouette scores for each layer using the rich library. The animation visualizes the separation between residual clouds across all layers as a GIF file.

Understanding the Research Implementation

Residual Extraction

The Model.get_residuals_batched() method in src/heretic/model.py computes hidden-state deltas between the original model and a null reference during forward passes. This yields tensors of shape (n_prompts, n_layers+1, hidden_dim) representing how each prompt affects the model's internal states.

Geometry Analysis Calculations

The Analyzer.print_residual_geometry() method performs several geometric computations per layer:

  1. Geometric Median: Computes the center of residual clusters using geom_median.torch.compute_geometric_median, which is more robust to outliers than arithmetic mean.
  2. Similarity Metrics: Calculates cosine similarities S(g,b), S(g,r), and S(b,r) using torch.nn.functional.cosine_similarity.
  3. Cluster Quality: Computes silhouette scores via sklearn.metrics.silhouette_score to quantify how well good and bad residuals separate.

PaCMAP Animation Generation

The Analyzer.plot_residuals() method creates the visualization through a multi-step process:

  1. Dimensionality Reduction: Applies PaCMAP (pacmap.PaCMAP) to project high-dimensional residuals to 2D, reusing each layer's projection as initialization for the next to ensure smooth transitions.
  2. Alignment: Rotates the embedding so the line connecting good and bad medians is horizontal, making layer-to-layer comparisons intuitive.
  3. Frame Interpolation: Generates N_TRANSITION_FRAMES = 20 linear interpolation frames between layers for fluid animation.
  4. Assembly: Uses imageio.v3.imwrite() to compile PNG sequences into animation.gif at the path specified by settings.residual_plot_path (default: plots/).

In src/heretic/main.py lines 445-450, the CLI automatically invokes these methods when the corresponding settings flags are enabled:

if settings.print_residual_geometry:
    analyzer.print_residual_geometry()

if settings.plot_residuals:
    analyzer.plot_residuals()

Summary

  • Install research dependencies with pip install heretic-llm[research]
  • Enable features via print_residual_geometry and plot_residuals in Settings
  • Use Analyzer class with residual tensors from Model.get_residuals_batched()
  • Outputs include a statistical geometry table and a PaCMAP animation GIF
  • Implementation relies on geometric median, scikit-learn, and PaCMAP libraries

Frequently Asked Questions

What Python packages are required for Heretic's research features?

The research features require geom-median, scikit-learn, rich, PaCMAP, imageio, matplotlib, and numpy. These are installed via the heretic-llm[research] extra. If missing, the analyzer methods in src/heretic/analyzer.py will raise an ImportError with installation instructions.

How do I interpret the residual geometry table output?

The table displays cosine similarities between good (harmless) and bad (harmful) residual vectors, L2-norms measuring vector magnitudes, and silhouette scores indicating cluster separation quality per layer. Higher silhouette scores indicate better separation between the refusal and compliance directions.

Can I customize the animation output path?

Yes, the output directory is controlled by the residual_plot_path field in the Settings class. The default location is plots/, but you can specify any path when constructing your Settings instance or in the TOML configuration file.

Why does Heretic use geometric median instead of mean for residual analysis?

The geometric median, computed via geom_median.torch.compute_geometric_median, provides robustness against outliers such as extremely toxic responses that might skew arithmetic means. This ensures the central tendency of residual clusters accurately represents typical prompt behavior.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →