How to Implement Model Quantization Using bnb_4bit with Heretic: A Complete Guide
You can implement model quantization using bnb_4bit with Heretic by setting quantization = "bnb_4bit" in your configuration file or using QuantizationMethod.BNB_4BIT in the Python API, which automatically generates a BitsAndBytesConfig and passes it to the Hugging Face model loader.
Heretic is an open-source framework designed for model abliteration and evaluation workflows. Implementing model quantization using bnb_4bit with Heretic enables you to load large language models in 4-bit precision, dramatically reducing GPU memory requirements through the bitsandbytes library integration.
Understanding bnb_4bit Quantization in Heretic
Heretic's architecture treats quantization as a first-class configuration option rather than an afterthought. The implementation spans configuration definitions, model loading logic, and special handling for adapter merging.
The QuantizationMethod Enum
The available quantization strategies are defined in src/heretic/config.py within the QuantizationMethod enum (lines 17-20). This enum includes BNB_4BIT as a valid option alongside NONE and other potential methods. When you specify quantization = "bnb_4bit" in your TOML configuration, Heretic parses this string into the corresponding enum member.
BitsAndBytesConfig Generation
The core quantization logic resides in src/heretic/model.py. The _get_quantization_config() method (lines 197-210) conditionally constructs a BitsAndBytesConfig object only when the selected method is BNB_4BIT. This configuration object specifies load_in_4bit=True and sets the compute dtype (typically bfloat16 or float16) for the dequantized computations.
Configuration Methods for bnb_4bit
Heretic provides two primary interfaces for enabling 4-bit quantization: declarative configuration files and programmatic Python API access.
TOML Configuration File
The simplest approach involves editing your config.toml file. Heretic automatically loads configuration from this file if present in the working directory.
# config.toml
quantization = "bnb_4bit" # Enable 4-bit quantization
device_map = "auto" # Automatic device placement
batch_size = 0 # Auto-tune batch size
After saving this configuration, run Heretic from the command line:
heretic Qwen/Qwen3-4B-Instruct-2507
The model loads with 4-bit precision, and you will see the confirmation message: [green]Ok[/] (quantized to 4-bit precision).
Python API Settings
For programmatic control, instantiate the Settings class with the QuantizationMethod.BNB_4BIT enum value.
from heretic import Settings, QuantizationMethod, Heretic
# Configure settings with bnb_4bit quantization
settings = Settings(
model="Qwen/Qwen3-4B-Instruct-2507",
quantization=QuantizationMethod.BNB_4BIT,
)
# Initialize Heretic - this automatically creates the BitsAndBytesConfig
heretic = Heretic(settings)
# Proceed with abliteration or evaluation workflows
heretic.run()
Loading Quantized Models
The model loading sequence in src/heretic/model.py (lines 101-112) demonstrates how the quantization configuration integrates with the Hugging Face ecosystem. The from_pretrained() method receives the quantization_config keyword argument only when _get_quantization_config() returns a valid BitsAndBytesConfig object. If the configuration specifies NONE, this argument remains None, and the model loads in full precision.
This conditional passing ensures that the same code path handles both quantized and non-quantized models without branching logic cluttering the main loading routine.
Special Handling for LoRA Adapters
Quantized models require specific handling when merging Low-Rank Adaptation (LoRA) weights. The 4-bit format cannot directly absorb LoRA weight updates due to precision constraints.
In src/heretic/model.py (lines 226-266), Heretic implements a merge workflow that:
- Reloads the base model in full precision (dequantized)
- Applies the LoRA adapter weights to the full-precision model
- Merges the adapter into the base weights
- Saves the merged model (now in full precision, which can be re-quantized if needed)
This approach ensures that quantization does not prevent you from using LoRA fine-tuning, though it requires temporary memory overhead during the merge operation.
Verification and Runtime Behavior
After successfully loading a quantized model, Heretic prints a confirmation message defined in src/heretic/model.py (lines 136-140):
[green]Ok[/] (quantized to 4-bit precision)
This output confirms that the BitsAndBytesConfig was correctly applied and the model tensors are stored in 4-bit format with compute dtype handling. You can verify memory savings by monitoring GPU VRAM usage compared to full-precision loading of the same architecture.
Summary
- Configuration: Set
quantization = "bnb_4bit"inconfig.tomlor useQuantizationMethod.BNB_4BITin Python to enable 4-bit quantization. - Implementation: Heretic automatically generates a
BitsAndBytesConfiginsrc/heretic/model.pyand passes it tofrom_pretrained(). - LoRA Handling: Merging adapters requires reloading the model in full precision temporarily, as implemented in lines 226-266 of the model loader.
- Verification: Successful quantization displays the confirmation message "(quantized to 4-bit precision)" during model initialization.
Frequently Asked Questions
What is bnb_4bit quantization?
bnb_4bit refers to 4-bit quantization implemented through the bitsandbytes library, which compresses model weights to 4-bit precision while maintaining computation in higher precision (typically float16 or bfloat16). This reduces GPU memory usage by approximately 75% compared to full 16-bit precision, enabling larger models to fit on consumer hardware.
Do I need to install bitsandbytes separately?
Yes, you must install the bitsandbytes package in your Python environment before using bnb_4bit quantization in Heretic. The library is not bundled with Heretic's core dependencies because it requires specific CUDA toolkit versions and platform-specific binaries. Install it via pip install bitsandbytes and verify GPU compatibility with the library's documentation.
Can I merge LoRA weights with a quantized model?
Yes, but with a specific workflow. Heretic handles LoRA merging with quantized models by first reloading the base model in full precision (dequantizing), then applying and merging the adapter weights, as implemented in src/heretic/model.py lines 226-266. You cannot merge LoRA weights directly into 4-bit tensors due to precision constraints, so this process requires temporary additional GPU memory during the merge operation.
How do I verify that quantization is active?
Check the console output during model loading. When Heretic successfully loads a model with bnb_4bit quantization, it prints the confirmation message [green]Ok[/] (quantized to 4-bit precision) as defined in src/heretic/model.py lines 136-140. Additionally, you can monitor GPU memory usage—quantized models should consume approximately 75% less VRAM than their full-precision counterparts.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →