What Model Architectures Does Heretic Support? Dense LLMs and Four MoE Families Explained

Heretic supports any dense transformer architecture out-of-the-box and specifically recognizes four distinct Mixture-of-Experts (MoE) families—Qwen 3, Phi-3.5-MoE, and two Granite MoE Hybrid variants—while currently excluding state space models (SSMs), hybrid layers, and novel attention mechanisms.

Heretic is a Python library designed for model abliteration and introspection of large language models. Understanding what model architectures Heretic supports is essential for researchers working with modern transformer variants beyond standard dense implementations. While the library universally handles dense transformers, its MoE compatibility targets specific implementations found in recent open-weight models.

Dense Transformer Support

Heretic works automatically with any dense transformer, including most multimodal variants that follow standard transformer layer conventions. According to the repository README at lines 66-68, the library is architected to handle dense models without requiring manual configuration or architecture-specific code paths. This universal compatibility applies to popular implementations like Llama, Mistral, and Qwen dense variants.

The detection logic resides in src/heretic/model.py, where the get_layer_modules method introspects model layers to identify standard MLP structures versus expert-routed alternatives.

MoE Architectures Supported by Heretic

Beyond dense transformers, Heretic explicitly implements detection and handling for four specific MoE families. The get_layer_modules method in src/heretic/model.py (lines 351-368) contains hardcoded detection patterns for each variant.

Qwen 3 MoE Detection

Heretic identifies Qwen 3 models (e.g., Qwen3-4B-Instruct-2507) by searching for an mlp.experts iterable inside each transformer layer. When this pattern is detected at lines 351-355, the library extracts individual expert parameters from the iterable container for abliteration analysis.

Phi-3.5-MoE Detection

For Phi-3.5-MoE and architecturally similar models, Heretic checks for a block_sparse_moe.experts container at lines 357-360. The detection logic specifically extracts each expert’s w2 weight matrix to enable per-expert introspection during the abliteration process.

Granite MoE Hybrid: Shared MLP Attention

Heretic handles the first variant of Granite MoE Hybrid architecture by detecting a shared_mlp.output_linear module within the attention branch (lines 361-364). This pattern identifies models that employ shared MLP layers alongside expert-routed components.

Granite MoE Hybrid: Expert MoE Layers

The second Granite variant is detected by finding a moe.experts list at lines 365-368. Heretic pulls the output_linear parameter from each expert in this list, enabling fine-grained analysis of the expert-specific components distinct from the shared attention mechanisms.

How Heretic Detects Architecture Types

The automatic architecture detection occurs during model initialization through the get_layer_modules method in src/heretic/model.py. This method sequentially checks for the attribute patterns listed above, returning the appropriate module collection for either dense MLPs or specific MoE expert groupings.

No manual configuration is required. When you load a model via Model(settings), Heretic introspects the PyTorch module hierarchy and automatically routes to the correct handler based on the presence of mlp.experts, block_sparse_moe.experts, or Granite-specific shared_mlp and moe attributes.

Usage Examples

Running Dense Models (Llama 3.1)

The following example demonstrates Heretic running on a standard dense transformer without requiring architecture-specific flags:

from heretic import Model, Settings, Prompt

settings = Settings(
    model="meta-llama/Meta-Llama-3.1-8B-Instruct",
    max_response_length=128,
    device_map="auto",
)
heretic = Model(settings)

# Simple test generation

response = heretic.get_responses([
    Prompt(system="You are a helpful assistant.", user="What is the capital of France?")
])
print(response[0])

Running MoE Models (Qwen 3)

For MoE models, Heretic automatically applies the appropriate detection logic. This example includes optional 4-bit quantization to accommodate larger models on limited VRAM:

from heretic import Model, Settings, Prompt

settings = Settings(
    model="Qwen/Qwen3-4B-Instruct-2507",
    max_response_length=128,
    device_map="auto",
    # optional: enable 4-bit quantization to fit on modest VRAM

    quantization="bnb_4bit",
)
heretic = Model(settings)

# Perform a quick generation to trigger MoE-specific handling

response = heretic.get_responses([
    Prompt(system="You are a helpful assistant.", user="Explain quantum tunnelling in simple terms.")
])
print(response[0])

Both snippets trigger the automatic module discovery in get_layer_modules, ensuring the correct abliteration logic is applied regardless of whether the underlying architecture is dense or MoE.

Unsupported Architectures

As documented in the README at lines 66-68, Heretic explicitly does not support state space models (SSMs), hybrid layers with non-standard topologies, inhomogeneous transformer layers, or architectures implementing novel attention mechanisms distinct from standard multi-head attention. Attempting to use Heretic with such models (e.g., Mamba, RetNet, or custom experimental hybrids) will likely result in incorrect layer identification or failed abliteration.

Summary

  • Heretic universally supports dense transformers including multimodal variants without manual configuration.
  • Four specific MoE families are explicitly handled: Qwen 3, Phi-3.5-MoE, and two Granite MoE Hybrid configurations (shared MLP attention and expert MoE layers).
  • Detection occurs automatically in src/heretic/model.py via get_layer_modules by inspecting layer attributes like mlp.experts and block_sparse_moe.experts at initialization.
  • State space models (SSMs) and other exotic architectures are currently unsupported according to README lines 66-68.
  • The library supports quantized MoE inference through the Settings configuration to reduce VRAM requirements.

Frequently Asked Questions

Does Heretic support state space models like Mamba or RetNet?

No. According to the README at lines 66-68, Heretic does not yet handle state space models (SSMs), hybrid layers, inhomogeneous layers, or novel attention mechanisms. The library focuses exclusively on dense transformers and specific MoE implementations with standard attention mechanisms.

How does Heretic automatically detect which MoE architecture a model uses?

Heretic inspects the module structure within each transformer layer through the get_layer_modules method in src/heretic/model.py. It checks for specific attribute patterns: mlp.experts indicates Qwen 3, block_sparse_moe.experts indicates Phi-3.5, and shared_mlp.output_linear or moe.experts indicate Granite variants. This detection requires no user configuration.

Can I use Heretic with quantized MoE models?

Yes. Heretic supports quantization options such as bnb_4bit through its Settings configuration, allowing MoE models like Qwen 3 to run on modest hardware. The quantization is applied during model loading in src/heretic/model.py before MoE-specific module discovery occurs.

Will Heretic work with custom or experimental MoE implementations?

Unlikely. Heretic only recognizes the four specific MoE families hardcoded in src/heretic/model.py lines 351-368. Custom MoE architectures with different module naming conventions, alternative expert routing mechanisms, or non-standard layer structures will not trigger the specialized MoE handlers and may be treated incorrectly as dense models.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →