How to Convert Hugging Face Safetensors to GGUF for BitNet: Complete Guide

BitNet provides a two-stage pipeline that converts Hugging Face safetensors checkpoints to GGUF format using gpu/convert_safetensors.py to reassemble sharded weights into PyTorch format, followed by utils/convert-hf-to-gguf-bitnet.py to generate the final quantized GGUF file.

The microsoft/BitNet repository includes specialized utilities designed to bridge the gap between Hugging Face's safetensors format and the GGUF format required for efficient BitNet inference. This guide explains the complete workflow to convert Hugging Face safetensors to GGUF for BitNet, covering both the automatic detection capabilities and the optional manual reassembly process.

Understanding the Two-Stage Conversion Pipeline

BitNet's conversion process is architected as a two-stage pipeline that handles the memory-mapped nature of safetensors while maintaining compatibility with the standard GGUF generation workflow.

Stage 1: Reassembling Safetensors into PyTorch Format

The first stage addresses the safetensors format's memory-mapped, read-only structure. In gpu/convert_safetensors.py, the convert_back function (lines 49-74) uses safetensors.safe_open to read tensors without fully deserializing them into Python objects, significantly reducing RAM usage during conversion.

This script specifically handles sharded checkpoints by building a TensorMap to track which tensors belong to which shards. It then concatenates split weights—particularly for Q, K, and V projection matrices—into complete tensors before writing a unified PyTorch state-dict (.pt) that matches the layout expected by the GGUF converter.

Stage 2: Converting to GGUF Format

The second stage utilizes utils/convert-hf-to-gguf-bitnet.py, which serves as the core Hugging Face to GGUF conversion engine. This script automatically detects whether the model directory contains .safetensors files via the Model._is_model_safetensors() method.

When safetensors are detected, the get_tensors method (lines 58-90) switches to using safetensors.safe_open instead of torch.load. The remainder of the conversion process—including tensor renaming, optional TL1/TL2 quantization, and vocabulary handling—proceeds identically to the standard .bin workflow, ensuring consistent GGUF output regardless of the original checkpoint format.

Step-by-Step Conversion Workflow

1. Install BitNet Conversion Utilities

Begin by cloning the repository and installing the required Python packages:

git clone https://github.com/microsoft/BitNet.git
cd BitNet
pip install -r requirements.txt

The requirements include safetensors, torch, gguf-py, numpy, and einops.

2. (Optional) Convert Safetensors to PyTorch Checkpoint

If you require an intermediate PyTorch state-dict for inspection or debugging, use the dedicated safetensors converter:

python gpu/convert_safetensors.py \
  --safetensors_file path/to/model.safetensors \
  --output model.pt \
  --model_name 2B

Select a --model_name that matches your model configuration (e.g., 2B). This produces a model.pt file containing the reassembled weights.

3. Run the GGUF Converter

Execute the main conversion script, pointing it to your model directory containing the safetensors files:

python utils/convert-hf-to-gguf-bitnet.py \
  --model path/to/model_dir \
  --outtype f16 \
  --outfile model.gguf

Available --outtype options include f32 (full precision), f16 (half precision), tl1 (ternary 1-bit), and tl2 (ternary 2-bit). The script automatically detects .safetensors files and processes them through the same pipeline as PyTorch bins.

4. Verify the Output

Confirm the conversion succeeded by inspecting the GGUF metadata:

gguf-tool info model.gguf

Alternatively, verify programmatically:

import gguf

gguf_file = gguf.GGUFReader('model.gguf')
print('Arch:', gguf_file.metadata['model_arch'])
print('Layers:', gguf_file.metadata['block_count'])
print('Quantisation:', gguf_file.metadata['file_type'])

Code Examples

Direct Conversion from a Safetensors Directory


# Assume the HF repo was cloned to ./my-bitnet-model

# The model directory contains model.safetensors + config.json etc.

python utils/convert-hf-to-gguf-bitnet.py \
    --model ./my-bitnet-model \
    --outtype tl2 \
    --outfile ./my-bitnet-model/bitnet.gguf

Using an Intermediate PyTorch Checkpoint


# Step 1: safetensors → torch checkpoint

python gpu/convert_safetensors.py \
    --safetensors_file ./my-bitnet-model/model.safetensors \
    --output ./my-bitnet-model/model.pt \
    --model_name 2B

# Step 2: torch checkpoint → GGUF

python utils/convert-hf-to-gguf-bitnet.py \
    --model ./my-bitnet-model \
    --outtype f16 \
    --outfile ./my-bitnet-model/bitnet-f16.gguf

Inspecting the Generated GGUF File

import gguf

gguf_file = gguf.GGUFReader('bitnet.gguf')
print('Architecture:', gguf_file.metadata['model_arch'])
print('Block count:', gguf_file.metadata['block_count'])
print('File type:', gguf_file.metadata['file_type'])

Key Implementation Files

The conversion pipeline relies on these specific files in the microsoft/BitNet repository:

File Role Key Components
gpu/convert_safetensors.py Reads .safetensors, reassembles sharded model parts, writes PyTorch checkpoint. convert_back function (lines 49-74), TensorMap construction
utils/convert-hf-to-gguf-bitnet.py Core HF-to-GGUF conversion engine; auto-detects safetensors and handles quantization. Model._is_model_safetensors(), get_tensors method (lines 58-90)
utils/convert.py Helper utilities shared across model families for vocabulary handling and tensor permutations. LlamaHfVocab, permute functions
gpu/gguf-py/gguf.py Vendored GGUF reader/writer implementation used throughout the conversion process. GGUFReader, GGUFWriter classes
gpu/convert_checkpoint.py Alternative checkpoint conversion entry point using the same Model abstraction. Checkpoint loading and validation logic

Summary

  • BitNet implements a two-stage conversion pipeline to transform Hugging Face safetensors into GGUF format, ensuring compatibility with efficient inference runtimes.
  • The gpu/convert_safetensors.py script handles memory-efficient loading of sharded safetensors using safe_open, reassembling Q, K, and V weights into a unified PyTorch state-dict.
  • utils/convert-hf-to-gguf-bitnet.py automatically detects safetensors files via _is_model_safetensors() and processes them through the same quantization pipeline (supporting f32, f16, tl1, and tl2 output types) used for standard PyTorch bins.
  • Supporting utilities in utils/convert.py and gpu/gguf-py/gguf.py provide vocabulary handling and low-level GGUF format operations.

Frequently Asked Questions

Can I convert safetensors directly to GGUF without creating an intermediate PyTorch file?

Yes. The utils/convert-hf-to-gguf-bitnet.py script automatically detects .safetensors files in your model directory using the Model._is_model_safetensors() method. When detected, it reads tensors directly using safetensors.safe_open within the get_tensors method, bypassing the need for an intermediate .pt file unless you specifically need one for debugging or manual inspection.

What quantization formats does BitNet support for GGUF output?

BitNet supports four primary output types controlled by the --outtype parameter: f32 (32-bit float), f16 (16-bit float), tl1 (ternary 1-bit), and tl2 (ternary 2-bit). The TL1 and TL2 formats are specific to BitNet's ternary quantization scheme, while f16 provides a good balance between precision and file size for standard inference.

How does BitNet handle sharded safetensors checkpoints?

The gpu/convert_safetensors.py script specifically handles sharded models through its convert_back function (lines 49-74). It constructs a TensorMap to track tensors across shards, then concatenates split weights—particularly for attention mechanisms where Q, K, and V matrices may be distributed across multiple files—into complete tensors before writing the unified checkpoint.

Why does BitNet use safetensors.safe_open instead of standard torch.load?

The conversion pipeline uses safetensors.safe_open to leverage memory-mapped file access, which allows reading tensor data without fully deserializing everything into Python objects. This approach significantly reduces RAM usage when processing large language models, prevents arbitrary code execution risks associated with pickle files, and enables efficient handling of sharded checkpoints by mapping only the required tensor segments into memory.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →