How to Add Support for New Model Architectures in Heretic: A Complete Guide

To add support for new model architectures in Heretic, you must modify three functions in src/heretic/model.py: get_model_class to select the correct AutoModel class, get_layers to expose the transformer layer stack, and get_layer_modules to identify the attention and MLP projection weights for abliteration.

Heretic is an open-source tool for "abliterating" harmful capabilities from large language models using LoRA-based weight editing. If you are working with the p-e-w/heretic repository and need to add support for new model architectures—whether encoder-decoder models like T5, vision-language models, or custom transformer layouts—you only need to adapt three specific functions in the model abstraction layer.

Step 1: Register the New Model Class in get_model_class

Heretic determines which Hugging Face AutoModel class to instantiate by inspecting the pretrained configuration. The get_model_class function (lines 36-44 in src/heretic/model.py) currently checks for vision_config to distinguish between multimodal and causal language models.

Detecting Architecture Types from Config

Add a new conditional branch to detect your architecture's unique configuration signature. For example, to support sequence-to-sequence models:

from transformers import AutoModelForSeq2SeqLM  # NEW import

def get_model_class(
    model: str,
) -> Type[AutoModelForImageTextToText] | Type[AutoModelForCausalLM] | Type[AutoModelForSeq2SeqLM]:
    configs = PretrainedConfig.get_config_dict(model)

    # Existing vision model check

    if any([("vision_config" in config) for config in configs]):
        return AutoModelForImageTextToText
    
    # NEW: Check for encoder-decoder architecture

    if any([("decoder" in config) for config in configs]):
        return AutoModelForSeq2SeqLM
    
    return AutoModelForCausalLM

This ensures Model.__init__ can call get_model_class(settings.model).from_pretrained(...) without raising a ClassNotFoundError.

Step 2: Locate Transformer Layers in get_layers

Once the model loads, Heretic must access the list of transformer layers to apply edits. The get_layers method (lines 12-25 in src/heretic/model.py) handles unwrapping PEFT adapters and navigating common layer hierarchies.

Handling Different Layer Hierarchies

If your architecture stores layers in a non-standard location (e.g., model.encoder.block instead of model.model.layers), add a new fallback:

def get_layers(self) -> ModuleList:
    model = self.model

    # Unwrap PEFT wrapper if present

    if isinstance(model, PeftModel):
        model = model.base_model.model

    # Existing: Multimodal path (e.g., Llava)

    with suppress(Exception):
        return model.model.language_model.layers

    # Existing: Text-only path (e.g., Llama, Mistral)

    with suppress(Exception):
        return model.model.layers

    # NEW: Encoder-decoder or encoder-only layouts (e.g., T5)

    with suppress(Exception):
        return model.encoder.block

    raise AssertionError("Unable to locate transformer layers for the loaded model.")

This method is invoked by the optimizer, analyzer, and residual-geometry utilities, so returning the correct ModuleList is critical for the rest of the pipeline to function.

Step 3: Map Attention and MLP Modules in get_layer_modules

Abliteration specifically targets the attention output projection (attn.o_proj) and MLP down-projection (mlp.down_proj) for LoRA-based editing. The get_layer_modules function (lines 26-74 in src/heretic/model.py) discovers these sub-modules using a helper called try_add.

Identifying Projection Layers for Abliteration

Different architectures use varying attribute names for these projections. Extend the discovery logic with additional try_add calls wrapped in suppress blocks:

def get_layer_modules(self, layer_index: int) -> dict[str, list[Module]]:
    layer = self.get_layers()[layer_index]
    modules: dict[str, list[Module]] = {}

    def try_add(component: str, module: Any):
        if isinstance(module, Module):
            modules.setdefault(component, []).append(module)
        else:
            assert not isinstance(module, Tensor), (
                f"Unexpected Tensor in {component} - expected nn.Module"
            )

    # Standard attention output projection (Llama, Mistral, etc.)

    try_add("attn.o_proj", layer.self_attn.o_proj)

    # Standard MLP down projection

    with suppress(Exception):
        try_add("mlp.down_proj", layer.mlp.down_proj)

    # ---------- NEW ARCHITECTURE EXAMPLES ----------

    
    # Example 1: Alternative attention naming (e.g., output_proj instead of o_proj)

    with suppress(Exception):
        try_add("attn.o_proj", layer.self_attn.output_proj)

    # Example 2: T5-style DenseReluDense blocks

    with suppress(Exception):
        try_add("mlp.down_proj", layer.DenseReluDense.wo)

    # Example 3: Custom expert modules in MoE architectures

    with suppress(Exception):
        for expert in getattr(layer, "custom_experts", []):
            try_add("mlp.down_proj", expert.down_proj)

    # -----------------------------------------------------------------

    
    total_modules = sum(len(v) for v in modules.values())
    assert total_modules > 0, "No abliterable modules found in layer"
    return modules

The abliterate method iterates over self.get_layer_modules(layer_index).items(), so adding these mappings enables Heretic to perform weight edits on the new architecture.

Complete Working Example: Adding T5 Support

Here is a consolidated example showing how to add support for T5-base, an encoder-decoder model:


# src/heretic/model.py

from transformers import AutoModelForSeq2SeqLM  # NEW import

def get_model_class(model: str):
    configs = PretrainedConfig.get_config_dict(model)
    
    if any([("vision_config" in cfg) for cfg in configs]):
        return AutoModelForImageTextToText
    if any([("decoder" in cfg) for cfg in configs]):  # NEW

        return AutoModelForSeq2SeqLM
    return AutoModelForCausalLM

def get_layers(self) -> ModuleList:
    model = self.model
    if isinstance(model, PeftModel):
        model = model.base_model.model
    
    with suppress(Exception):
        return model.model.language_model.layers
    with suppress(Exception):
        return model.model.layers
    with suppress(Exception):
        return model.encoder.block  # NEW: T5 encoder blocks

    
    raise AssertionError("Unable to locate transformer layers")

def get_layer_modules(self, layer_index: int):
    layer = self.get_layers()[layer_index]
    modules = {}
    
    def try_add(component: str, module: Any):
        if isinstance(module, Module):
            modules.setdefault(component, []).append(module)
    
    # Standard attempts

    with suppress(Exception):
        try_add("attn.o_proj", layer.self_attn.o_proj)
    with suppress(Exception):
        try_add("mlp.down_proj", layer.mlp.down_proj)
    
    # T5-specific mappings

    with suppress(Exception):
        try_add("attn.o_proj", layer.SelfAttention.o)  # NEW

    with suppress(Exception):
        try_add("mlp.down_proj", layer.DenseReluDense.wo)  # NEW

    
    assert sum(len(v) for v in modules.values()) > 0
    return modules

With these changes, you can run Heretic with --model t5-base and the optimizer will successfully load the model, enumerate its encoder layers, and apply LoRA-based abliteration to the attention and MLP projections.

Key Files and Functions Reference

File Responsibility Key Functions to Modify
src/heretic/model.py Model loading and layer abstraction get_model_class (lines 36-44), get_layers (lines 12-25), get_layer_modules (lines 26-74)
src/heretic/config.py Configuration schema Only modify if adding new architecture-specific settings
src/heretic/utils.py Helper utilities Generally unchanged for new architectures

Summary

  • Register the model class by extending get_model_class in src/heretic/model.py to return the appropriate AutoModel... class based on configuration keys.
  • Expose the layer stack by adding fallback logic to get_layers to handle non-standard layer hierarchies like model.encoder.block.
  • Map the projections by adding try_add calls in get_layer_modules to identify attention output and MLP down-projection weights regardless of naming conventions.
  • Test with a real model such as T5 to verify that Heretic can load, analyze, and abliterate the new architecture.

Frequently Asked Questions

Do I need to modify the abliteration logic itself to support new architectures?

No. The abliteration algorithms in Heretic are architecture-agnostic and operate on the generic attn.o_proj and mlp.down_proj components returned by get_layer_modules. As long as you correctly map these projections in src/heretic/model.py, the existing LoRA-based editing and residual analysis code will work without modification.

How do I handle models with different attention mechanisms like multi-query attention?

Multi-query attention (MQA) and grouped-query attention (GQA) typically reuse the same o_proj (output projection) attribute names as standard multi-head attention. You should verify the exact attribute path in your model's Hugging Face implementation, then add a try_add("attn.o_proj", layer.self_attn.o_proj) variant (or layer.self_attn.q_proj if targeting queries) wrapped in a suppress(Exception) block to handle the specific naming convention.

Can I add support for quantized models like GPTQ or AWQ?

Yes, but you typically do not need to modify src/heretic/model.py for quantization support. Heretic uses transformers.AutoModel... classes which automatically handle quantization config via from_pretrained. Ensure your quantization config is passed during model loading in the configuration settings. The layer enumeration and module discovery logic remains the same because quantization wrappers preserve the underlying attribute structure.

What if my model uses a custom HuggingFace architecture not in the standard library?

For custom architectures, you must ensure the model class is importable in your Python environment (e.g., via trust_remote_code=True). Then follow the same three-step process: register the custom class in get_model_class, add the layer path in get_layers, and map the projection attributes in get_layer_modules. You may need to inspect the custom model's modeling_*.py file to find the exact attribute names for attention and MLP components.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →