# Difference Between Merging a Model and Using an Adapter with Heretic

> Understand merging models vs using adapters in Heretic. Merge bakes LoRAs into base weights for standalone checkpoints. Adapters keep LoRAs separate for lightweight experimentation. Learn the key differences now.

- Repository: [Philipp Emanuel Weidmann/heretic](https://github.com/p-e-w/heretic)
- Tags: deep-dive
- Published: 2026-02-19

---

**Merging a model bakes LoRA adapters permanently into the base weights to create a standalone checkpoint, while using an adapter keeps the small LoRA matrices separate, enabling lightweight storage and rapid experimentation without modifying the original model.**

Heretic is a tool for experimenting with language models using **PEFT LoRA adapters** to modify base models without touching the original weights. When working with Heretic, you must decide whether to keep adapters external or merge them into the base model. This choice affects storage costs, inference behavior, and your ability to iterate quickly.

## What Is the Difference Between Merging and Adapters in Heretic?

Heretic implements two distinct strategies for handling LoRA adaptations. The **adapter-only** strategy maintains modularity, while the **merge** strategy creates a self-contained model.

### The Adapter-Only Strategy

When you use the adapter strategy, Heretic attaches LoRA modules to the base model at load-time via `Model._apply_lora()` (lines 63‑95 of [`src/heretic/model.py`](https://github.com/p-e-w/heretic/blob/main/src/heretic/model.py)). Only the small LoRA weight matrices (`lora_A` and `lora_B`) are saved to disk, while the original model files remain untouched. This approach is implemented in the adapter-save branch of [`src/heretic/main.py`](https://github.com/p-e-w/heretic/blob/main/src/heretic/main.py) (lines 60‑66).

### The Merge Strategy

The merge strategy calls `Model.get_merged_model()` (lines 22‑66 of [`src/heretic/model.py`](https://github.com/p-e-w/heretic/blob/main/src/heretic/model.py)) to permanently combine the adapter weights with a full-precision copy of the base model. The method invokes `merge_and_unload()` to bake the adaptations into the base tensors, then removes the LoRA modules entirely. This creates a new `PreTrainedModel` checkpoint that contains no PEFT components. The merge-save branch in [`main.py`](https://github.com/p-e-w/heretic/blob/main/main.py) (lines 64‑67) handles the persistence of this merged checkpoint.

| Aspect | Adapter-Only | Merged Model |
|--------|--------------|--------------|
| **Storage** | Only LoRA matrices (~few MB) | Full model checkpoint (hundreds of MB/GB) |
| **Base Model** | Unchanged on disk | Permanently modified |
| **Memory at Runtime** | Base model + small adapter tensors | Single model with baked-in weights |
| **Flexibility** | Switch adapters instantly without reload | Must reload base model to change adaptations |
| **4-bit Handling** | Works natively | Requires CPU offloading and full-precision copy (lines 27‑58) |

## How Heretic Implements LoRA Adapters

Heretic creates LoRA adapters dynamically when you instantiate the `Model` class. The `_apply_lora()` method inspects the target modules and injects trainable low-rank matrices alongside the frozen base weights.

```python
from heretic.model import Model
from heretic.config import Settings

settings = Settings(
    model="meta-llama/Meta-Llama-3-8B",
    dtypes=["bfloat16"],
    quantization="none",
    row_normalization="full"
)

# LoRA adapters are applied automatically via _apply_lora()

model = Model(settings)

```

This approach keeps the original checkpoint intact while allowing you to train or modify only the adapter parameters.

## How to Merge Adapters into a Base Model

When you need a self-contained artifact for production or for environments without PEFT support, use the merge strategy. Heretic’s `get_merged_model()` handles the complexity of merging, including special logic for quantized models.

```python

# Obtain the merged model (permanently bakes in adapters)

merged = model.get_merged_model()

# Save as a standard Hugging Face checkpoint

merged.save_pretrained("./merged_model")
model.tokenizer.save_pretrained("./merged_model")

```

### Special Handling for 4-Bit Quantization

For BitsAndBytes 4-bit quantized models, Heretic cannot merge adapters directly on the GPU because the base weights are compressed. Instead, `get_merged_model()` (lines 27‑58 of [`model.py`](https://github.com/p-e-w/heretic/blob/main/model.py)) reloads the base model in full precision on the CPU, copies the adapter weights, performs the merge, and then moves the result to the target device. This avoids VRAM overflow while still producing a merged checkpoint.

## When to Use Each Strategy

Choose the **adapter-only** strategy when you need flexibility and efficient storage. This is ideal for rapid prototyping, A/B testing different adaptation techniques, or when you want to distribute only the small adapter files while requiring users to possess the base model separately.

Choose the **merge** strategy when you need a single deployable artifact. This is necessary for production inference pipelines that do not support PEFT, for exporting to formats that require standard model checkpoints, or when you want to eliminate the small runtime overhead of adapter computation.

## Summary

- **Adapter-only** keeps LoRA matrices separate from the base model, storing only the small `lora_A` and `lora_B` weights via `model.model.save_pretrained()`.
- **Merging** permanently combines adapters with base weights using `get_merged_model()` and `merge_and_unload()`, creating a standalone `PreTrainedModel`.
- Adapters enable fast resets and low storage costs, while merged models provide self-contained checkpoints suitable for non-PEFT environments.
- For 4-bit quantized models, Heretic handles merging by temporarily reloading the base model in full precision on CPU to avoid memory issues.

## Frequently Asked Questions

### Does merging a model reduce inference speed?

Merging removes the overhead of computing adapter updates during the forward pass, which can slightly improve inference speed compared to the adapter-only approach. However, the primary benefit is compatibility with standard inference pipelines that do not support PEFT, rather than significant performance gains.

### Can I unmerge a model after merging?

No, once you merge adapters into the base model using `get_merged_model()`, the LoRA modules are removed and the changes are baked into the base weight tensors. To recover the separate components, you must reload the original base model and reattach the adapter files if you preserved them separately.

### How much disk space do adapters save compared to merged models?

Adapters typically require only a few megabytes of storage because they consist of small low-rank matrices, whereas merged models require the full checkpoint size—often hundreds of megabytes to several gigabytes depending on the base model. This makes adapters ideal for version control and distribution.

### Does Heretic support 4-bit quantization with adapters?

Yes, Heretic fully supports using LoRA adapters with 4-bit quantized models via BitsAndBytes. When using the adapter-only strategy, the base model remains quantized while adapters stay in full precision. When merging, Heretic automatically handles the complexity by reloading the base model in full precision on CPU before merging to avoid VRAM overflow.