How to Implement Model Quantization Using bnb_4bit with Heretic: A Complete Guide

You can implement model quantization using bnb_4bit with Heretic by setting quantization = "bnb_4bit" in your configuration file or using QuantizationMethod.BNB_4BIT in the Python API, which automatically generates a BitsAndBytesConfig and passes it to the Hugging Face model loader.

Heretic is an open-source framework designed for model abliteration and evaluation workflows. Implementing model quantization using bnb_4bit with Heretic enables you to load large language models in 4-bit precision, dramatically reducing GPU memory requirements through the bitsandbytes library integration.

Understanding bnb_4bit Quantization in Heretic

Heretic's architecture treats quantization as a first-class configuration option rather than an afterthought. The implementation spans configuration definitions, model loading logic, and special handling for adapter merging.

The QuantizationMethod Enum

The available quantization strategies are defined in src/heretic/config.py within the QuantizationMethod enum (lines 17-20). This enum includes BNB_4BIT as a valid option alongside NONE and other potential methods. When you specify quantization = "bnb_4bit" in your TOML configuration, Heretic parses this string into the corresponding enum member.

BitsAndBytesConfig Generation

The core quantization logic resides in src/heretic/model.py. The _get_quantization_config() method (lines 197-210) conditionally constructs a BitsAndBytesConfig object only when the selected method is BNB_4BIT. This configuration object specifies load_in_4bit=True and sets the compute dtype (typically bfloat16 or float16) for the dequantized computations.

Configuration Methods for bnb_4bit

Heretic provides two primary interfaces for enabling 4-bit quantization: declarative configuration files and programmatic Python API access.

TOML Configuration File

The simplest approach involves editing your config.toml file. Heretic automatically loads configuration from this file if present in the working directory.


# config.toml

quantization = "bnb_4bit"          # Enable 4-bit quantization

device_map   = "auto"              # Automatic device placement

batch_size   = 0                   # Auto-tune batch size

After saving this configuration, run Heretic from the command line:

heretic Qwen/Qwen3-4B-Instruct-2507

The model loads with 4-bit precision, and you will see the confirmation message: [green]Ok[/] (quantized to 4-bit precision).

Python API Settings

For programmatic control, instantiate the Settings class with the QuantizationMethod.BNB_4BIT enum value.

from heretic import Settings, QuantizationMethod, Heretic

# Configure settings with bnb_4bit quantization

settings = Settings(
    model="Qwen/Qwen3-4B-Instruct-2507",
    quantization=QuantizationMethod.BNB_4BIT,
)

# Initialize Heretic - this automatically creates the BitsAndBytesConfig

heretic = Heretic(settings)

# Proceed with abliteration or evaluation workflows

heretic.run()

Loading Quantized Models

The model loading sequence in src/heretic/model.py (lines 101-112) demonstrates how the quantization configuration integrates with the Hugging Face ecosystem. The from_pretrained() method receives the quantization_config keyword argument only when _get_quantization_config() returns a valid BitsAndBytesConfig object. If the configuration specifies NONE, this argument remains None, and the model loads in full precision.

This conditional passing ensures that the same code path handles both quantized and non-quantized models without branching logic cluttering the main loading routine.

Special Handling for LoRA Adapters

Quantized models require specific handling when merging Low-Rank Adaptation (LoRA) weights. The 4-bit format cannot directly absorb LoRA weight updates due to precision constraints.

In src/heretic/model.py (lines 226-266), Heretic implements a merge workflow that:

  1. Reloads the base model in full precision (dequantized)
  2. Applies the LoRA adapter weights to the full-precision model
  3. Merges the adapter into the base weights
  4. Saves the merged model (now in full precision, which can be re-quantized if needed)

This approach ensures that quantization does not prevent you from using LoRA fine-tuning, though it requires temporary memory overhead during the merge operation.

Verification and Runtime Behavior

After successfully loading a quantized model, Heretic prints a confirmation message defined in src/heretic/model.py (lines 136-140):


[green]Ok[/] (quantized to 4-bit precision)

This output confirms that the BitsAndBytesConfig was correctly applied and the model tensors are stored in 4-bit format with compute dtype handling. You can verify memory savings by monitoring GPU VRAM usage compared to full-precision loading of the same architecture.

Summary

  • Configuration: Set quantization = "bnb_4bit" in config.toml or use QuantizationMethod.BNB_4BIT in Python to enable 4-bit quantization.
  • Implementation: Heretic automatically generates a BitsAndBytesConfig in src/heretic/model.py and passes it to from_pretrained().
  • LoRA Handling: Merging adapters requires reloading the model in full precision temporarily, as implemented in lines 226-266 of the model loader.
  • Verification: Successful quantization displays the confirmation message "(quantized to 4-bit precision)" during model initialization.

Frequently Asked Questions

What is bnb_4bit quantization?

bnb_4bit refers to 4-bit quantization implemented through the bitsandbytes library, which compresses model weights to 4-bit precision while maintaining computation in higher precision (typically float16 or bfloat16). This reduces GPU memory usage by approximately 75% compared to full 16-bit precision, enabling larger models to fit on consumer hardware.

Do I need to install bitsandbytes separately?

Yes, you must install the bitsandbytes package in your Python environment before using bnb_4bit quantization in Heretic. The library is not bundled with Heretic's core dependencies because it requires specific CUDA toolkit versions and platform-specific binaries. Install it via pip install bitsandbytes and verify GPU compatibility with the library's documentation.

Can I merge LoRA weights with a quantized model?

Yes, but with a specific workflow. Heretic handles LoRA merging with quantized models by first reloading the base model in full precision (dequantizing), then applying and merging the adapter weights, as implemented in src/heretic/model.py lines 226-266. You cannot merge LoRA weights directly into 4-bit tensors due to precision constraints, so this process requires temporary additional GPU memory during the merge operation.

How do I verify that quantization is active?

Check the console output during model loading. When Heretic successfully loads a model with bnb_4bit quantization, it prints the confirmation message [green]Ok[/] (quantized to 4-bit precision) as defined in src/heretic/model.py lines 136-140. Additionally, you can monitor GPU memory usage—quantized models should consume approximately 75% less VRAM than their full-precision counterparts.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →