Heretic Ablation Parameters: Understanding the Four Key Controls for Model Modification
The four key ablation parameters in Heretic—max_weight, max_weight_position, min_weight, and min_weight_distance—control how aggressively and where LoRA adapters modify transformer weights to remove unwanted behaviors.
Heretic implements parameter-driven ablative fine-tuning to modify transformer model weights through LoRA adapters, targeting specific refusal directions while preserving general capabilities. The core mechanism relies on the AbliterationParameters dataclass defined in src/heretic/model.py, which provides granular, spatially-aware control over how each layer is nudged away from undesired directions. Understanding these four parameters is essential for tuning the strength and scope of model modifications.
What Are the Key Ablation Parameters in Heretic?
The AbliterationParameters dataclass (lines 47-53 in src/heretic/model.py) defines four fields that determine how the ablative update is calculated for every layer component:
-
max_weight(float): The strongest (most negative) scaling factor applied to the layer closest to the target position. This serves as the base λ in the LoRA weight updatelora_B = -λ · vfor layers atmax_weight_position. -
max_weight_position(float): The layer index (possibly fractional) wheremax_weightshould be applied. This determines the reference layer for the ablation. -
min_weight(float): The weakest (least negative) scaling factor applied to layers farther from the target position. This serves as the endpoint of the linear interpolation. -
min_weight_distance(float): The radius (in layer units) over which the interpolation occurs. Layers beyond this distance frommax_weight_positionare skipped entirely, leaving their weights untouched.
These parameters allow you to create a spatial gradient of ablation strength, concentrating the strongest effect at a specific layer while tapering off toward surrounding layers.
How Ablation Parameters Work in the Abliteration Process
The Model.abliterate method consumes these parameters to compute per-layer weight factors through linear interpolation. For each layer/component pair, the method calculates the distance from the target position and determines whether to apply ablation:
# Retrieve the abliteration parameters for the current component
params = parameters[component]
distance = abs(layer_index - params.max_weight_position)
# Skip layers outside the effective radius
if distance > params.min_weight_distance:
continue
# Linear interpolation between max and min weight
weight = params.max_weight + (distance / params.min_weight_distance) * (
params.min_weight - params.max_weight
)
The interpolated weight then constructs the LoRA adapters that push the model’s weight matrix W away from the refusal direction v:
lora_A = (v @ W).view(1, -1) # vᵀ · W
lora_B = (-weight * v).view(-1, 1) # -λ · v
If row normalization is enabled (settings.row_normalization != RowNormalization.NONE), additional processing (norm preservation and optional low-rank SVD) occurs before writing the adapters back to the model (see lines 470-513 in src/heretic/model.py).
Practical Implementation: Configuring Ablation Parameters
Define Parameters for Specific Components
Target the MLP down-projection with aggressive ablation around layer 12, fading out over ±3 layers:
from heretic.model import AbliterationParameters
mlp_params = AbliterationParameters(
max_weight=-0.8, # strongest negative scaling
max_weight_position=12.0, # target layer index
min_weight=-0.1, # mild scaling far away
min_weight_distance=3.0, # affect layers 9-15
)
Assemble the Parameter Dictionary
Map components to their respective parameters using keys that match model.get_abliterable_components():
params = {
"mlp.down_proj": mlp_params,
# Additional components: "attn.o_proj", "mlp.up_proj", etc.
}
Execute the Ablation
Apply the configured parameters through the abliterate method:
from heretic.model import Model
from heretic.config import Settings
settings = Settings.from_toml_path("config.default.toml")
model = Model(settings)
# Example refusal direction (unit vector)
import torch
refusal_dir = torch.randn(model.get_layers()[0].self_attn.o_proj.weight.shape[0])
refusal_dir = torch.nn.functional.normalize(refusal_dir, p=2, dim=0)
model.abliterate(
refusal_directions=refusal_dir.unsqueeze(0),
direction_index=None,
parameters=params,
)
Verify the Modification
Test the ablated model's behavior:
from heretic.utils import Prompt
prompt = Prompt(system="You are a helpful assistant.", user="Explain quantum entanglement.")
response = model.get_responses([prompt])[0]
print(response)
Summary
- Four core parameters (
max_weight,max_weight_position,min_weight,min_weight_distance) insrc/heretic/model.pycontrol the spatial distribution and intensity of ablative modifications. - Linear interpolation between
max_weightandmin_weightcreates a tapered effect across layers based on distance frommax_weight_position. - LoRA adapters implement the actual weight updates via
lora_A = (v @ W).view(1, -1)andlora_B = (-weight * v).view(-1, 1). - Spatial boundaries defined by
min_weight_distanceensure only targeted sub-networks are altered, leaving distant layers untouched. - Row normalization settings in
src/heretic/config.pyprovide additional control over weight matrix properties during ablation.
Frequently Asked Questions
How do I choose the right values for max_weight and min_weight?
Start with max_weight between -0.5 and -1.0 for strong ablation, and min_weight between -0.1 and 0.0 for gentle tapering. According to the Heretic source code, these values represent negative scaling factors in the LoRA update equation, so more negative values push the model further from the refusal direction. The optimal values depend on how deeply embedded the target behavior is in specific layers.
Can I apply different ablation parameters to different model components?
Yes, the parameters argument accepts a dictionary mapping component names to AbliterationParameters instances. As implemented in src/heretic/model.py, you can specify distinct parameters for "mlp.down_proj", "attn.o_proj", "mlp.up_proj", and other components returned by model.get_abliterable_components(), allowing fine-grained control over which parts of the transformer are modified.
What happens if min_weight_distance is set to zero?
Setting min_weight_distance to zero causes the ablation to skip all layers entirely. The method checks if distance > params.min_weight_distance before applying updates, so a zero value means only layers at exactly the max_weight_position (distance zero) would qualify, but the interpolation formula would still apply. In practice, use values ≥1.0 to ensure meaningful spatial coverage.
How does max_weight_position handle fractional layer indices?
The parameter accepts float values, allowing you to position the peak ablation effect between discrete transformer layers. The code calculates distance = abs(layer_index - params.max_weight_position), so a value of 12.5 would place the strongest effect halfway between layers 12 and 13, with surrounding layers receiving proportionally interpolated weights based on their distance from this fractional position.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →