# How to Implement Model Quantization Using bnb_4bit with Heretic: A Complete Guide

> Implement model quantization using bnb_4bit with Heretic. This guide shows how to configure quantization via file or the Python API for efficient model loading with Hugging Face.

- Repository: [Philipp Emanuel Weidmann/heretic](https://github.com/p-e-w/heretic)
- Tags: how-to-guide
- Published: 2026-02-19

---

**You can implement model quantization using bnb_4bit with Heretic by setting `quantization = "bnb_4bit"` in your configuration file or using `QuantizationMethod.BNB_4BIT` in the Python API, which automatically generates a `BitsAndBytesConfig` and passes it to the Hugging Face model loader.**

Heretic is an open-source framework designed for model abliteration and evaluation workflows. Implementing model quantization using bnb_4bit with Heretic enables you to load large language models in 4-bit precision, dramatically reducing GPU memory requirements through the **bitsandbytes** library integration.

## Understanding bnb_4bit Quantization in Heretic

Heretic's architecture treats quantization as a first-class configuration option rather than an afterthought. The implementation spans configuration definitions, model loading logic, and special handling for adapter merging.

### The QuantizationMethod Enum

The available quantization strategies are defined in [`src/heretic/config.py`](https://github.com/p-e-w/heretic/blob/main/src/heretic/config.py) within the `QuantizationMethod` enum (lines 17-20). This enum includes `BNB_4BIT` as a valid option alongside `NONE` and other potential methods. When you specify `quantization = "bnb_4bit"` in your TOML configuration, Heretic parses this string into the corresponding enum member.

### BitsAndBytesConfig Generation

The core quantization logic resides in [`src/heretic/model.py`](https://github.com/p-e-w/heretic/blob/main/src/heretic/model.py). The `_get_quantization_config()` method (lines 197-210) conditionally constructs a `BitsAndBytesConfig` object only when the selected method is `BNB_4BIT`. This configuration object specifies `load_in_4bit=True` and sets the compute dtype (typically `bfloat16` or `float16`) for the dequantized computations.

## Configuration Methods for bnb_4bit

Heretic provides two primary interfaces for enabling 4-bit quantization: declarative configuration files and programmatic Python API access.

### TOML Configuration File

The simplest approach involves editing your [`config.toml`](https://github.com/p-e-w/heretic/blob/main/config.toml) file. Heretic automatically loads configuration from this file if present in the working directory.

```toml

# config.toml

quantization = "bnb_4bit"          # Enable 4-bit quantization

device_map   = "auto"              # Automatic device placement

batch_size   = 0                   # Auto-tune batch size

```

After saving this configuration, run Heretic from the command line:

```bash
heretic Qwen/Qwen3-4B-Instruct-2507

```

The model loads with 4-bit precision, and you will see the confirmation message: `[green]Ok[/] (quantized to 4-bit precision)`.

### Python API Settings

For programmatic control, instantiate the `Settings` class with the `QuantizationMethod.BNB_4BIT` enum value.

```python
from heretic import Settings, QuantizationMethod, Heretic

# Configure settings with bnb_4bit quantization

settings = Settings(
    model="Qwen/Qwen3-4B-Instruct-2507",
    quantization=QuantizationMethod.BNB_4BIT,
)

# Initialize Heretic - this automatically creates the BitsAndBytesConfig

heretic = Heretic(settings)

# Proceed with abliteration or evaluation workflows

heretic.run()

```

## Loading Quantized Models

The model loading sequence in [`src/heretic/model.py`](https://github.com/p-e-w/heretic/blob/main/src/heretic/model.py) (lines 101-112) demonstrates how the quantization configuration integrates with the Hugging Face ecosystem. The `from_pretrained()` method receives the `quantization_config` keyword argument only when `_get_quantization_config()` returns a valid `BitsAndBytesConfig` object. If the configuration specifies `NONE`, this argument remains `None`, and the model loads in full precision.

This conditional passing ensures that the same code path handles both quantized and non-quantized models without branching logic cluttering the main loading routine.

## Special Handling for LoRA Adapters

Quantized models require specific handling when merging Low-Rank Adaptation (LoRA) weights. The 4-bit format cannot directly absorb LoRA weight updates due to precision constraints.

In [`src/heretic/model.py`](https://github.com/p-e-w/heretic/blob/main/src/heretic/model.py) (lines 226-266), Heretic implements a merge workflow that:
1. Reloads the base model in full precision (dequantized)
2. Applies the LoRA adapter weights to the full-precision model
3. Merges the adapter into the base weights
4. Saves the merged model (now in full precision, which can be re-quantized if needed)

This approach ensures that quantization does not prevent you from using LoRA fine-tuning, though it requires temporary memory overhead during the merge operation.

## Verification and Runtime Behavior

After successfully loading a quantized model, Heretic prints a confirmation message defined in [`src/heretic/model.py`](https://github.com/p-e-w/heretic/blob/main/src/heretic/model.py) (lines 136-140):

```

[green]Ok[/] (quantized to 4-bit precision)

```

This output confirms that the `BitsAndBytesConfig` was correctly applied and the model tensors are stored in 4-bit format with compute dtype handling. You can verify memory savings by monitoring GPU VRAM usage compared to full-precision loading of the same architecture.

## Summary

- **Configuration**: Set `quantization = "bnb_4bit"` in [`config.toml`](https://github.com/p-e-w/heretic/blob/main/config.toml) or use `QuantizationMethod.BNB_4BIT` in Python to enable 4-bit quantization.
- **Implementation**: Heretic automatically generates a `BitsAndBytesConfig` in [`src/heretic/model.py`](https://github.com/p-e-w/heretic/blob/main/src/heretic/model.py) and passes it to `from_pretrained()`.
- **LoRA Handling**: Merging adapters requires reloading the model in full precision temporarily, as implemented in lines 226-266 of the model loader.
- **Verification**: Successful quantization displays the confirmation message "(quantized to 4-bit precision)" during model initialization.

## Frequently Asked Questions

### What is bnb_4bit quantization?

**bnb_4bit** refers to 4-bit quantization implemented through the **bitsandbytes** library, which compresses model weights to 4-bit precision while maintaining computation in higher precision (typically float16 or bfloat16). This reduces GPU memory usage by approximately 75% compared to full 16-bit precision, enabling larger models to fit on consumer hardware.

### Do I need to install bitsandbytes separately?

**Yes**, you must install the `bitsandbytes` package in your Python environment before using `bnb_4bit` quantization in Heretic. The library is not bundled with Heretic's core dependencies because it requires specific CUDA toolkit versions and platform-specific binaries. Install it via `pip install bitsandbytes` and verify GPU compatibility with the library's documentation.

### Can I merge LoRA weights with a quantized model?

**Yes, but with a specific workflow**. Heretic handles LoRA merging with quantized models by first reloading the base model in full precision (dequantizing), then applying and merging the adapter weights, as implemented in [`src/heretic/model.py`](https://github.com/p-e-w/heretic/blob/main/src/heretic/model.py) lines 226-266. You cannot merge LoRA weights directly into 4-bit tensors due to precision constraints, so this process requires temporary additional GPU memory during the merge operation.

### How do I verify that quantization is active?

**Check the console output during model loading**. When Heretic successfully loads a model with `bnb_4bit` quantization, it prints the confirmation message `[green]Ok[/] (quantized to 4-bit precision)` as defined in [`src/heretic/model.py`](https://github.com/p-e-w/heretic/blob/main/src/heretic/model.py) lines 136-140. Additionally, you can monitor GPU memory usage—quantized models should consume approximately 75% less VRAM than their full-precision counterparts.