BitNet Inference Preprocessing: Converting Hugging Face Checkpoints to I2S GGUF

BitNet inference requires a two-stage preprocessing pipeline that first quantizes projection matrices in Hugging Face safetensors checkpoints using utils/preprocess-huggingface-bitnet.py, then converts the result to an I2S GGUF file via utils/convert-helper-bitnet.py to enable fast 1-bit kernel execution.

Before running inference with microsoft/BitNet, raw Hugging Face checkpoints must undergo specific BitNet inference preprocessing to transform weights into the packed 2-bit representation required by the optimized CPU and GPU kernels. This guide explains the exact steps, source files, and commands needed to convert safetensors models into the ready-to-load GGUF format.

The Two-Stage BitNet Inference Preprocessing Pipeline

The microsoft/BitNet repository implements a strict two-stage workflow to prepare models for deployment. According to the source code in utils/convert-helper-bitnet.py (lines 68-96), the pipeline sequentially applies fp16-style quantization to projection weights, converts the checkpoint to GGUF format, and finally packs weights into the lossless I2S representation.

Stage 1: Quantize Projection Matrices with preprocess-huggingface-bitnet.py

The first stage processes the original .safetensors checkpoint to prepare projection matrices for 1-bit computation. In utils/preprocess-huggingface-bitnet.py (lines 28-31), the script identifies all linear projection layers—including q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, and down_proj—and applies the quant_weight_fp16 function to scale weights to the integer range -1 to 1.

This scaling is mandatory because BitNet's kernels require weights to be pre-quantized before packing into the 2-bit I2S format. The script overwrites the original safetensors file with quantized tensors, logging [INFO] Quantizing … for each matched tensor.

Execute this stage manually with:

python utils/preprocess-huggingface-bitnet.py \
    --input models/BitNet-b1.58-2B-4T/model.safetensors \
    --output models/BitNet-b1.58-2B-4T/model.safetensors

Stage 2: Convert to GGUF and Apply I2S Quantization

The second stage converts the pre-quantized safetensors into the GGUF container format and applies the final I2S quantization. The utils/convert-helper-bitnet.py script orchestrates three internal commands:

  1. Preprocessing: Runs preprocess-huggingface-bitnet.py (if not already done)
  2. F32 GGUF conversion: Executes utils/convert-ms-to-gguf-bitnet.py to generate ggml-model-f32-bitnet.gguf
  3. I2S quantization: Invokes the compiled llama-quantize binary (located in build/bin/) to produce the final ggml-model-i2s-bitnet.gguf

The I2S format represents the 1-bit weights in a packed, lossless structure that matches the kernel implementations in gpu/model.py. Run the full conversion with:

python utils/convert-helper-bitnet.py models/BitNet-b1.58-2B-4T

Optional: Quantizing Embedding Matrices

For additional memory bandwidth reduction, you can quantize the embedding matrix to fp16 before GGUF conversion. The setup_env.py script accepts the --quant-embd flag when configuring the model environment, as shown in the source argument parsing.

Enable embedding quantization during setup:

python setup_env.py -md models/BitNet-b1.58-2B-4T -q i2_s --quant-embd

Complete End-to-End Preprocessing Example

The following workflow demonstrates the complete BitNet inference preprocessing pipeline from repository cloning to final model generation:


# Clone and setup environment

git clone --recursive https://github.com/microsoft/BitNet.git
cd BitNet
conda create -n bitnet python=3.9 -y && conda activate bitnet
pip install -r requirements.txt

# Download BitNet model in safetensors format

huggingface-cli download microsoft/BitNet-b1.58-2B-4T --local-dir models/BitNet-b1.58-2B-4T

# Stage 1: Quantize projection matrices

python utils/preprocess-huggingface-bitnet.py \
    --input models/BitNet-b1.58-2B-4T/model.safetensors \
    --output models/BitNet-b1.58-2B-4T/model.safetensors

# Stage 2: Convert to GGUF and generate I2S quantized model

python utils/convert-helper-bitnet.py models/BitNet-b1.58-2B-4T

# Verify output

ls models/BitNet-b1.58-2B-4T/ggml-model-i2s-bitnet.gguf

Loading Preprocessed Models for Inference

Once preprocessing completes, the resulting ggml-model-i2s-bitnet.gguf file serves as the input for all BitNet inference engines. For CPU inference, use run_inference.py:

python run_inference.py \
    -m models/BitNet-b1.58-2B-4T/ggml-model-i2s-bitnet.gguf \
    -p "Once upon a time" -n 64

For GPU inference, ensure you have built the kernels with clang >= 18, then start the server:

./bitnet_server -m models/BitNet-b1.58-2B-4T/ggml-model-i2s-bitnet.gguf

Summary

  • BitNet inference preprocessing requires quantizing Hugging Face checkpoints through two distinct stages before they can be loaded by the runtime.
  • Stage 1 uses utils/preprocess-huggingface-bitnet.py to scale projection matrices (q_proj, k_proj, v_proj, etc.) to the -1...1 range using quant_weight_fp16.
  • Stage 2 employs utils/convert-helper-bitnet.py to generate an F32 GGUF intermediate, then converts it to the final I2S GGUF format via the llama-quantize binary.
  • The output ggml-model-i2s-bitnet.gguf file is the only format accepted by the optimized CPU (run_inference.py) and GPU (bitnet_server) inference engines.
  • Optional embedding quantization via setup_env.py --quant-embd reduces memory footprint without accuracy loss for most tasks.

Frequently Asked Questions

What file format does BitNet inference require?

BitNet inference engines require models in I2S GGUF format, specifically the ggml-model-i2s-bitnet.gguf file produced by the preprocessing pipeline. Raw Hugging Face safetensors checkpoints cannot be loaded directly; they must first be processed through the two-stage quantization and conversion workflow described in utils/convert-helper-bitnet.py.

Which projection matrices does the preprocessing script quantize?

The utils/preprocess-huggingface-bitnet.py script targets seven specific linear projection layers: q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, and down_proj. These matrices are quantized using the quant_weight_fp16 function to prepare them for the 1-bit kernel operations implemented in the BitNet runtime.

Can I skip the preprocessing and convert a standard GGUF model to I2S?

No. Standard GGUF models lack the fp16-scaled projection weights that BitNet kernels expect. You must first run preprocess-huggingface-bitnet.py on the original safetensors checkpoint to apply the specific quantization scaling before converting to GGUF and applying I2S packing. Attempting to quantize an unprepared model will result in runtime errors or incorrect outputs.

How do I enable embedding quantization during preprocessing?

Add the --quant-embd flag when running setup_env.py with the -q i2_s argument. This triggers additional fp16 quantization of the embedding matrix before GGUF conversion, reducing model size and memory bandwidth requirements. The command format is: python setup_env.py -md <model_path> -q i2_s --quant-embd.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →