Difference Between Merging a Model and Using an Adapter with Heretic
Merging a model bakes LoRA adapters permanently into the base weights to create a standalone checkpoint, while using an adapter keeps the small LoRA matrices separate, enabling lightweight storage and rapid experimentation without modifying the original model.
Heretic is a tool for experimenting with language models using PEFT LoRA adapters to modify base models without touching the original weights. When working with Heretic, you must decide whether to keep adapters external or merge them into the base model. This choice affects storage costs, inference behavior, and your ability to iterate quickly.
What Is the Difference Between Merging and Adapters in Heretic?
Heretic implements two distinct strategies for handling LoRA adaptations. The adapter-only strategy maintains modularity, while the merge strategy creates a self-contained model.
The Adapter-Only Strategy
When you use the adapter strategy, Heretic attaches LoRA modules to the base model at load-time via Model._apply_lora() (lines 63‑95 of src/heretic/model.py). Only the small LoRA weight matrices (lora_A and lora_B) are saved to disk, while the original model files remain untouched. This approach is implemented in the adapter-save branch of src/heretic/main.py (lines 60‑66).
The Merge Strategy
The merge strategy calls Model.get_merged_model() (lines 22‑66 of src/heretic/model.py) to permanently combine the adapter weights with a full-precision copy of the base model. The method invokes merge_and_unload() to bake the adaptations into the base tensors, then removes the LoRA modules entirely. This creates a new PreTrainedModel checkpoint that contains no PEFT components. The merge-save branch in main.py (lines 64‑67) handles the persistence of this merged checkpoint.
| Aspect | Adapter-Only | Merged Model |
|---|---|---|
| Storage | Only LoRA matrices (~few MB) | Full model checkpoint (hundreds of MB/GB) |
| Base Model | Unchanged on disk | Permanently modified |
| Memory at Runtime | Base model + small adapter tensors | Single model with baked-in weights |
| Flexibility | Switch adapters instantly without reload | Must reload base model to change adaptations |
| 4-bit Handling | Works natively | Requires CPU offloading and full-precision copy (lines 27‑58) |
How Heretic Implements LoRA Adapters
Heretic creates LoRA adapters dynamically when you instantiate the Model class. The _apply_lora() method inspects the target modules and injects trainable low-rank matrices alongside the frozen base weights.
from heretic.model import Model
from heretic.config import Settings
settings = Settings(
model="meta-llama/Meta-Llama-3-8B",
dtypes=["bfloat16"],
quantization="none",
row_normalization="full"
)
# LoRA adapters are applied automatically via _apply_lora()
model = Model(settings)
This approach keeps the original checkpoint intact while allowing you to train or modify only the adapter parameters.
How to Merge Adapters into a Base Model
When you need a self-contained artifact for production or for environments without PEFT support, use the merge strategy. Heretic’s get_merged_model() handles the complexity of merging, including special logic for quantized models.
# Obtain the merged model (permanently bakes in adapters)
merged = model.get_merged_model()
# Save as a standard Hugging Face checkpoint
merged.save_pretrained("./merged_model")
model.tokenizer.save_pretrained("./merged_model")
Special Handling for 4-Bit Quantization
For BitsAndBytes 4-bit quantized models, Heretic cannot merge adapters directly on the GPU because the base weights are compressed. Instead, get_merged_model() (lines 27‑58 of model.py) reloads the base model in full precision on the CPU, copies the adapter weights, performs the merge, and then moves the result to the target device. This avoids VRAM overflow while still producing a merged checkpoint.
When to Use Each Strategy
Choose the adapter-only strategy when you need flexibility and efficient storage. This is ideal for rapid prototyping, A/B testing different adaptation techniques, or when you want to distribute only the small adapter files while requiring users to possess the base model separately.
Choose the merge strategy when you need a single deployable artifact. This is necessary for production inference pipelines that do not support PEFT, for exporting to formats that require standard model checkpoints, or when you want to eliminate the small runtime overhead of adapter computation.
Summary
- Adapter-only keeps LoRA matrices separate from the base model, storing only the small
lora_Aandlora_Bweights viamodel.model.save_pretrained(). - Merging permanently combines adapters with base weights using
get_merged_model()andmerge_and_unload(), creating a standalonePreTrainedModel. - Adapters enable fast resets and low storage costs, while merged models provide self-contained checkpoints suitable for non-PEFT environments.
- For 4-bit quantized models, Heretic handles merging by temporarily reloading the base model in full precision on CPU to avoid memory issues.
Frequently Asked Questions
Does merging a model reduce inference speed?
Merging removes the overhead of computing adapter updates during the forward pass, which can slightly improve inference speed compared to the adapter-only approach. However, the primary benefit is compatibility with standard inference pipelines that do not support PEFT, rather than significant performance gains.
Can I unmerge a model after merging?
No, once you merge adapters into the base model using get_merged_model(), the LoRA modules are removed and the changes are baked into the base weight tensors. To recover the separate components, you must reload the original base model and reattach the adapter files if you preserved them separately.
How much disk space do adapters save compared to merged models?
Adapters typically require only a few megabytes of storage because they consist of small low-rank matrices, whereas merged models require the full checkpoint size—often hundreds of megabytes to several gigabytes depending on the base model. This makes adapters ideal for version control and distribution.
Does Heretic support 4-bit quantization with adapters?
Yes, Heretic fully supports using LoRA adapters with 4-bit quantized models via BitsAndBytes. When using the adapter-only strategy, the base model remains quantized while adapters stay in full precision. When merging, Heretic automatically handles the complexity by reloading the base model in full precision on CPU before merging to avoid VRAM overflow.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →