How to Add Support for New Model Architectures in Heretic: A Complete Guide
To add support for new model architectures in Heretic, you must modify three functions in src/heretic/model.py: get_model_class to select the correct AutoModel class, get_layers to expose the transformer layer stack, and get_layer_modules to identify the attention and MLP projection weights for abliteration.
Heretic is an open-source tool for "abliterating" harmful capabilities from large language models using LoRA-based weight editing. If you are working with the p-e-w/heretic repository and need to add support for new model architectures—whether encoder-decoder models like T5, vision-language models, or custom transformer layouts—you only need to adapt three specific functions in the model abstraction layer.
Step 1: Register the New Model Class in get_model_class
Heretic determines which Hugging Face AutoModel class to instantiate by inspecting the pretrained configuration. The get_model_class function (lines 36-44 in src/heretic/model.py) currently checks for vision_config to distinguish between multimodal and causal language models.
Detecting Architecture Types from Config
Add a new conditional branch to detect your architecture's unique configuration signature. For example, to support sequence-to-sequence models:
from transformers import AutoModelForSeq2SeqLM # NEW import
def get_model_class(
model: str,
) -> Type[AutoModelForImageTextToText] | Type[AutoModelForCausalLM] | Type[AutoModelForSeq2SeqLM]:
configs = PretrainedConfig.get_config_dict(model)
# Existing vision model check
if any([("vision_config" in config) for config in configs]):
return AutoModelForImageTextToText
# NEW: Check for encoder-decoder architecture
if any([("decoder" in config) for config in configs]):
return AutoModelForSeq2SeqLM
return AutoModelForCausalLM
This ensures Model.__init__ can call get_model_class(settings.model).from_pretrained(...) without raising a ClassNotFoundError.
Step 2: Locate Transformer Layers in get_layers
Once the model loads, Heretic must access the list of transformer layers to apply edits. The get_layers method (lines 12-25 in src/heretic/model.py) handles unwrapping PEFT adapters and navigating common layer hierarchies.
Handling Different Layer Hierarchies
If your architecture stores layers in a non-standard location (e.g., model.encoder.block instead of model.model.layers), add a new fallback:
def get_layers(self) -> ModuleList:
model = self.model
# Unwrap PEFT wrapper if present
if isinstance(model, PeftModel):
model = model.base_model.model
# Existing: Multimodal path (e.g., Llava)
with suppress(Exception):
return model.model.language_model.layers
# Existing: Text-only path (e.g., Llama, Mistral)
with suppress(Exception):
return model.model.layers
# NEW: Encoder-decoder or encoder-only layouts (e.g., T5)
with suppress(Exception):
return model.encoder.block
raise AssertionError("Unable to locate transformer layers for the loaded model.")
This method is invoked by the optimizer, analyzer, and residual-geometry utilities, so returning the correct ModuleList is critical for the rest of the pipeline to function.
Step 3: Map Attention and MLP Modules in get_layer_modules
Abliteration specifically targets the attention output projection (attn.o_proj) and MLP down-projection (mlp.down_proj) for LoRA-based editing. The get_layer_modules function (lines 26-74 in src/heretic/model.py) discovers these sub-modules using a helper called try_add.
Identifying Projection Layers for Abliteration
Different architectures use varying attribute names for these projections. Extend the discovery logic with additional try_add calls wrapped in suppress blocks:
def get_layer_modules(self, layer_index: int) -> dict[str, list[Module]]:
layer = self.get_layers()[layer_index]
modules: dict[str, list[Module]] = {}
def try_add(component: str, module: Any):
if isinstance(module, Module):
modules.setdefault(component, []).append(module)
else:
assert not isinstance(module, Tensor), (
f"Unexpected Tensor in {component} - expected nn.Module"
)
# Standard attention output projection (Llama, Mistral, etc.)
try_add("attn.o_proj", layer.self_attn.o_proj)
# Standard MLP down projection
with suppress(Exception):
try_add("mlp.down_proj", layer.mlp.down_proj)
# ---------- NEW ARCHITECTURE EXAMPLES ----------
# Example 1: Alternative attention naming (e.g., output_proj instead of o_proj)
with suppress(Exception):
try_add("attn.o_proj", layer.self_attn.output_proj)
# Example 2: T5-style DenseReluDense blocks
with suppress(Exception):
try_add("mlp.down_proj", layer.DenseReluDense.wo)
# Example 3: Custom expert modules in MoE architectures
with suppress(Exception):
for expert in getattr(layer, "custom_experts", []):
try_add("mlp.down_proj", expert.down_proj)
# -----------------------------------------------------------------
total_modules = sum(len(v) for v in modules.values())
assert total_modules > 0, "No abliterable modules found in layer"
return modules
The abliterate method iterates over self.get_layer_modules(layer_index).items(), so adding these mappings enables Heretic to perform weight edits on the new architecture.
Complete Working Example: Adding T5 Support
Here is a consolidated example showing how to add support for T5-base, an encoder-decoder model:
# src/heretic/model.py
from transformers import AutoModelForSeq2SeqLM # NEW import
def get_model_class(model: str):
configs = PretrainedConfig.get_config_dict(model)
if any([("vision_config" in cfg) for cfg in configs]):
return AutoModelForImageTextToText
if any([("decoder" in cfg) for cfg in configs]): # NEW
return AutoModelForSeq2SeqLM
return AutoModelForCausalLM
def get_layers(self) -> ModuleList:
model = self.model
if isinstance(model, PeftModel):
model = model.base_model.model
with suppress(Exception):
return model.model.language_model.layers
with suppress(Exception):
return model.model.layers
with suppress(Exception):
return model.encoder.block # NEW: T5 encoder blocks
raise AssertionError("Unable to locate transformer layers")
def get_layer_modules(self, layer_index: int):
layer = self.get_layers()[layer_index]
modules = {}
def try_add(component: str, module: Any):
if isinstance(module, Module):
modules.setdefault(component, []).append(module)
# Standard attempts
with suppress(Exception):
try_add("attn.o_proj", layer.self_attn.o_proj)
with suppress(Exception):
try_add("mlp.down_proj", layer.mlp.down_proj)
# T5-specific mappings
with suppress(Exception):
try_add("attn.o_proj", layer.SelfAttention.o) # NEW
with suppress(Exception):
try_add("mlp.down_proj", layer.DenseReluDense.wo) # NEW
assert sum(len(v) for v in modules.values()) > 0
return modules
With these changes, you can run Heretic with --model t5-base and the optimizer will successfully load the model, enumerate its encoder layers, and apply LoRA-based abliteration to the attention and MLP projections.
Key Files and Functions Reference
| File | Responsibility | Key Functions to Modify |
|---|---|---|
src/heretic/model.py |
Model loading and layer abstraction | get_model_class (lines 36-44), get_layers (lines 12-25), get_layer_modules (lines 26-74) |
src/heretic/config.py |
Configuration schema | Only modify if adding new architecture-specific settings |
src/heretic/utils.py |
Helper utilities | Generally unchanged for new architectures |
Summary
- Register the model class by extending
get_model_classinsrc/heretic/model.pyto return the appropriateAutoModel...class based on configuration keys. - Expose the layer stack by adding fallback logic to
get_layersto handle non-standard layer hierarchies likemodel.encoder.block. - Map the projections by adding
try_addcalls inget_layer_modulesto identify attention output and MLP down-projection weights regardless of naming conventions. - Test with a real model such as T5 to verify that Heretic can load, analyze, and abliterate the new architecture.
Frequently Asked Questions
Do I need to modify the abliteration logic itself to support new architectures?
No. The abliteration algorithms in Heretic are architecture-agnostic and operate on the generic attn.o_proj and mlp.down_proj components returned by get_layer_modules. As long as you correctly map these projections in src/heretic/model.py, the existing LoRA-based editing and residual analysis code will work without modification.
How do I handle models with different attention mechanisms like multi-query attention?
Multi-query attention (MQA) and grouped-query attention (GQA) typically reuse the same o_proj (output projection) attribute names as standard multi-head attention. You should verify the exact attribute path in your model's Hugging Face implementation, then add a try_add("attn.o_proj", layer.self_attn.o_proj) variant (or layer.self_attn.q_proj if targeting queries) wrapped in a suppress(Exception) block to handle the specific naming convention.
Can I add support for quantized models like GPTQ or AWQ?
Yes, but you typically do not need to modify src/heretic/model.py for quantization support. Heretic uses transformers.AutoModel... classes which automatically handle quantization config via from_pretrained. Ensure your quantization config is passed during model loading in the configuration settings. The layer enumeration and module discovery logic remains the same because quantization wrappers preserve the underlying attribute structure.
What if my model uses a custom HuggingFace architecture not in the standard library?
For custom architectures, you must ensure the model class is importable in your Python environment (e.g., via trust_remote_code=True). Then follow the same three-step process: register the custom class in get_model_class, add the layer path in get_layers, and map the projection attributes in get_layer_modules. You may need to inspect the custom model's modeling_*.py file to find the exact attribute names for attention and MLP components.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →