# BitNet Inference Preprocessing: Converting Hugging Face Checkpoints to I2S GGUF

> Learn the two-stage preprocessing for BitNet inference: quantize Hugging Face checkpoints and convert to I2S GGUF for fast 1-bit kernel execution. Follow our guide now.

- Repository: [Microsoft/BitNet](https://github.com/microsoft/BitNet)
- Tags: how-to-guide
- Published: 2026-03-13

---

**BitNet inference requires a two-stage preprocessing pipeline that first quantizes projection matrices in Hugging Face safetensors checkpoints using [`utils/preprocess-huggingface-bitnet.py`](https://github.com/microsoft/BitNet/blob/main/utils/preprocess-huggingface-bitnet.py), then converts the result to an I2S GGUF file via [`utils/convert-helper-bitnet.py`](https://github.com/microsoft/BitNet/blob/main/utils/convert-helper-bitnet.py) to enable fast 1-bit kernel execution.**

Before running inference with microsoft/BitNet, raw Hugging Face checkpoints must undergo specific **BitNet inference preprocessing** to transform weights into the packed 2-bit representation required by the optimized CPU and GPU kernels. This guide explains the exact steps, source files, and commands needed to convert safetensors models into the ready-to-load GGUF format.

## The Two-Stage BitNet Inference Preprocessing Pipeline

The microsoft/BitNet repository implements a strict two-stage workflow to prepare models for deployment. According to the source code in [`utils/convert-helper-bitnet.py`](https://github.com/microsoft/BitNet/blob/main/utils/convert-helper-bitnet.py) (lines 68-96), the pipeline sequentially applies fp16-style quantization to projection weights, converts the checkpoint to GGUF format, and finally packs weights into the lossless I2S representation.

### Stage 1: Quantize Projection Matrices with preprocess-huggingface-bitnet.py

The first stage processes the original `.safetensors` checkpoint to prepare projection matrices for 1-bit computation. In [`utils/preprocess-huggingface-bitnet.py`](https://github.com/microsoft/BitNet/blob/main/utils/preprocess-huggingface-bitnet.py) (lines 28-31), the script identifies all linear projection layers—including `q_proj`, `k_proj`, `v_proj`, `o_proj`, `gate_proj`, `up_proj`, and `down_proj`—and applies the `quant_weight_fp16` function to scale weights to the integer range -1 to 1.

This scaling is mandatory because BitNet's kernels require weights to be pre-quantized before packing into the 2-bit I2S format. The script overwrites the original safetensors file with quantized tensors, logging `[INFO] Quantizing …` for each matched tensor.

Execute this stage manually with:

```bash
python utils/preprocess-huggingface-bitnet.py \
    --input models/BitNet-b1.58-2B-4T/model.safetensors \
    --output models/BitNet-b1.58-2B-4T/model.safetensors

```

### Stage 2: Convert to GGUF and Apply I2S Quantization

The second stage converts the pre-quantized safetensors into the GGUF container format and applies the final I2S quantization. The [`utils/convert-helper-bitnet.py`](https://github.com/microsoft/BitNet/blob/main/utils/convert-helper-bitnet.py) script orchestrates three internal commands:

1. **Preprocessing**: Runs [`preprocess-huggingface-bitnet.py`](https://github.com/microsoft/BitNet/blob/main/preprocess-huggingface-bitnet.py) (if not already done)
2. **F32 GGUF conversion**: Executes [`utils/convert-ms-to-gguf-bitnet.py`](https://github.com/microsoft/BitNet/blob/main/utils/convert-ms-to-gguf-bitnet.py) to generate `ggml-model-f32-bitnet.gguf`
3. **I2S quantization**: Invokes the compiled `llama-quantize` binary (located in `build/bin/`) to produce the final `ggml-model-i2s-bitnet.gguf`

The I2S format represents the 1-bit weights in a packed, lossless structure that matches the kernel implementations in [`gpu/model.py`](https://github.com/microsoft/BitNet/blob/main/gpu/model.py). Run the full conversion with:

```bash
python utils/convert-helper-bitnet.py models/BitNet-b1.58-2B-4T

```

## Optional: Quantizing Embedding Matrices

For additional memory bandwidth reduction, you can quantize the embedding matrix to fp16 before GGUF conversion. The [`setup_env.py`](https://github.com/microsoft/BitNet/blob/main/setup_env.py) script accepts the `--quant-embd` flag when configuring the model environment, as shown in the source argument parsing.

Enable embedding quantization during setup:

```bash
python setup_env.py -md models/BitNet-b1.58-2B-4T -q i2_s --quant-embd

```

## Complete End-to-End Preprocessing Example

The following workflow demonstrates the complete **BitNet inference preprocessing** pipeline from repository cloning to final model generation:

```bash

# Clone and setup environment

git clone --recursive https://github.com/microsoft/BitNet.git
cd BitNet
conda create -n bitnet python=3.9 -y && conda activate bitnet
pip install -r requirements.txt

# Download BitNet model in safetensors format

huggingface-cli download microsoft/BitNet-b1.58-2B-4T --local-dir models/BitNet-b1.58-2B-4T

# Stage 1: Quantize projection matrices

python utils/preprocess-huggingface-bitnet.py \
    --input models/BitNet-b1.58-2B-4T/model.safetensors \
    --output models/BitNet-b1.58-2B-4T/model.safetensors

# Stage 2: Convert to GGUF and generate I2S quantized model

python utils/convert-helper-bitnet.py models/BitNet-b1.58-2B-4T

# Verify output

ls models/BitNet-b1.58-2B-4T/ggml-model-i2s-bitnet.gguf

```

## Loading Preprocessed Models for Inference

Once preprocessing completes, the resulting `ggml-model-i2s-bitnet.gguf` file serves as the input for all BitNet inference engines. For CPU inference, use [`run_inference.py`](https://github.com/microsoft/BitNet/blob/main/run_inference.py):

```bash
python run_inference.py \
    -m models/BitNet-b1.58-2B-4T/ggml-model-i2s-bitnet.gguf \
    -p "Once upon a time" -n 64

```

For GPU inference, ensure you have built the kernels with clang >= 18, then start the server:

```bash
./bitnet_server -m models/BitNet-b1.58-2B-4T/ggml-model-i2s-bitnet.gguf

```

## Summary

- **BitNet inference preprocessing** requires quantizing Hugging Face checkpoints through two distinct stages before they can be loaded by the runtime.
- **Stage 1** uses [`utils/preprocess-huggingface-bitnet.py`](https://github.com/microsoft/BitNet/blob/main/utils/preprocess-huggingface-bitnet.py) to scale projection matrices (`q_proj`, `k_proj`, `v_proj`, etc.) to the -1...1 range using `quant_weight_fp16`.
- **Stage 2** employs [`utils/convert-helper-bitnet.py`](https://github.com/microsoft/BitNet/blob/main/utils/convert-helper-bitnet.py) to generate an F32 GGUF intermediate, then converts it to the final I2S GGUF format via the `llama-quantize` binary.
- The output `ggml-model-i2s-bitnet.gguf` file is the only format accepted by the optimized CPU ([`run_inference.py`](https://github.com/microsoft/BitNet/blob/main/run_inference.py)) and GPU (`bitnet_server`) inference engines.
- Optional embedding quantization via `setup_env.py --quant-embd` reduces memory footprint without accuracy loss for most tasks.

## Frequently Asked Questions

### What file format does BitNet inference require?

BitNet inference engines require models in **I2S GGUF format**, specifically the `ggml-model-i2s-bitnet.gguf` file produced by the preprocessing pipeline. Raw Hugging Face safetensors checkpoints cannot be loaded directly; they must first be processed through the two-stage quantization and conversion workflow described in [`utils/convert-helper-bitnet.py`](https://github.com/microsoft/BitNet/blob/main/utils/convert-helper-bitnet.py).

### Which projection matrices does the preprocessing script quantize?

The [`utils/preprocess-huggingface-bitnet.py`](https://github.com/microsoft/BitNet/blob/main/utils/preprocess-huggingface-bitnet.py) script targets seven specific linear projection layers: `q_proj`, `k_proj`, `v_proj`, `o_proj`, `gate_proj`, `up_proj`, and `down_proj`. These matrices are quantized using the `quant_weight_fp16` function to prepare them for the 1-bit kernel operations implemented in the BitNet runtime.

### Can I skip the preprocessing and convert a standard GGUF model to I2S?

No. Standard GGUF models lack the fp16-scaled projection weights that BitNet kernels expect. You must first run [`preprocess-huggingface-bitnet.py`](https://github.com/microsoft/BitNet/blob/main/preprocess-huggingface-bitnet.py) on the original safetensors checkpoint to apply the specific quantization scaling before converting to GGUF and applying I2S packing. Attempting to quantize an unprepared model will result in runtime errors or incorrect outputs.

### How do I enable embedding quantization during preprocessing?

Add the `--quant-embd` flag when running [`setup_env.py`](https://github.com/microsoft/BitNet/blob/main/setup_env.py) with the `-q i2_s` argument. This triggers additional fp16 quantization of the embedding matrix before GGUF conversion, reducing model size and memory bandwidth requirements. The command format is: `python setup_env.py -md <model_path> -q i2_s --quant-embd`.