How to Convert Hugging Face Safetensors to GGUF for BitNet: Complete Guide
BitNet provides a two-stage pipeline that converts Hugging Face safetensors checkpoints to GGUF format using gpu/convert_safetensors.py to reassemble sharded weights into PyTorch format, followed by utils/convert-hf-to-gguf-bitnet.py to generate the final quantized GGUF file.
The microsoft/BitNet repository includes specialized utilities designed to bridge the gap between Hugging Face's safetensors format and the GGUF format required for efficient BitNet inference. This guide explains the complete workflow to convert Hugging Face safetensors to GGUF for BitNet, covering both the automatic detection capabilities and the optional manual reassembly process.
Understanding the Two-Stage Conversion Pipeline
BitNet's conversion process is architected as a two-stage pipeline that handles the memory-mapped nature of safetensors while maintaining compatibility with the standard GGUF generation workflow.
Stage 1: Reassembling Safetensors into PyTorch Format
The first stage addresses the safetensors format's memory-mapped, read-only structure. In gpu/convert_safetensors.py, the convert_back function (lines 49-74) uses safetensors.safe_open to read tensors without fully deserializing them into Python objects, significantly reducing RAM usage during conversion.
This script specifically handles sharded checkpoints by building a TensorMap to track which tensors belong to which shards. It then concatenates split weights—particularly for Q, K, and V projection matrices—into complete tensors before writing a unified PyTorch state-dict (.pt) that matches the layout expected by the GGUF converter.
Stage 2: Converting to GGUF Format
The second stage utilizes utils/convert-hf-to-gguf-bitnet.py, which serves as the core Hugging Face to GGUF conversion engine. This script automatically detects whether the model directory contains .safetensors files via the Model._is_model_safetensors() method.
When safetensors are detected, the get_tensors method (lines 58-90) switches to using safetensors.safe_open instead of torch.load. The remainder of the conversion process—including tensor renaming, optional TL1/TL2 quantization, and vocabulary handling—proceeds identically to the standard .bin workflow, ensuring consistent GGUF output regardless of the original checkpoint format.
Step-by-Step Conversion Workflow
1. Install BitNet Conversion Utilities
Begin by cloning the repository and installing the required Python packages:
git clone https://github.com/microsoft/BitNet.git
cd BitNet
pip install -r requirements.txt
The requirements include safetensors, torch, gguf-py, numpy, and einops.
2. (Optional) Convert Safetensors to PyTorch Checkpoint
If you require an intermediate PyTorch state-dict for inspection or debugging, use the dedicated safetensors converter:
python gpu/convert_safetensors.py \
--safetensors_file path/to/model.safetensors \
--output model.pt \
--model_name 2B
Select a --model_name that matches your model configuration (e.g., 2B). This produces a model.pt file containing the reassembled weights.
3. Run the GGUF Converter
Execute the main conversion script, pointing it to your model directory containing the safetensors files:
python utils/convert-hf-to-gguf-bitnet.py \
--model path/to/model_dir \
--outtype f16 \
--outfile model.gguf
Available --outtype options include f32 (full precision), f16 (half precision), tl1 (ternary 1-bit), and tl2 (ternary 2-bit). The script automatically detects .safetensors files and processes them through the same pipeline as PyTorch bins.
4. Verify the Output
Confirm the conversion succeeded by inspecting the GGUF metadata:
gguf-tool info model.gguf
Alternatively, verify programmatically:
import gguf
gguf_file = gguf.GGUFReader('model.gguf')
print('Arch:', gguf_file.metadata['model_arch'])
print('Layers:', gguf_file.metadata['block_count'])
print('Quantisation:', gguf_file.metadata['file_type'])
Code Examples
Direct Conversion from a Safetensors Directory
# Assume the HF repo was cloned to ./my-bitnet-model
# The model directory contains model.safetensors + config.json etc.
python utils/convert-hf-to-gguf-bitnet.py \
--model ./my-bitnet-model \
--outtype tl2 \
--outfile ./my-bitnet-model/bitnet.gguf
Using an Intermediate PyTorch Checkpoint
# Step 1: safetensors → torch checkpoint
python gpu/convert_safetensors.py \
--safetensors_file ./my-bitnet-model/model.safetensors \
--output ./my-bitnet-model/model.pt \
--model_name 2B
# Step 2: torch checkpoint → GGUF
python utils/convert-hf-to-gguf-bitnet.py \
--model ./my-bitnet-model \
--outtype f16 \
--outfile ./my-bitnet-model/bitnet-f16.gguf
Inspecting the Generated GGUF File
import gguf
gguf_file = gguf.GGUFReader('bitnet.gguf')
print('Architecture:', gguf_file.metadata['model_arch'])
print('Block count:', gguf_file.metadata['block_count'])
print('File type:', gguf_file.metadata['file_type'])
Key Implementation Files
The conversion pipeline relies on these specific files in the microsoft/BitNet repository:
| File | Role | Key Components |
|---|---|---|
gpu/convert_safetensors.py |
Reads .safetensors, reassembles sharded model parts, writes PyTorch checkpoint. |
convert_back function (lines 49-74), TensorMap construction |
utils/convert-hf-to-gguf-bitnet.py |
Core HF-to-GGUF conversion engine; auto-detects safetensors and handles quantization. | Model._is_model_safetensors(), get_tensors method (lines 58-90) |
utils/convert.py |
Helper utilities shared across model families for vocabulary handling and tensor permutations. | LlamaHfVocab, permute functions |
gpu/gguf-py/gguf.py |
Vendored GGUF reader/writer implementation used throughout the conversion process. | GGUFReader, GGUFWriter classes |
gpu/convert_checkpoint.py |
Alternative checkpoint conversion entry point using the same Model abstraction. | Checkpoint loading and validation logic |
Summary
- BitNet implements a two-stage conversion pipeline to transform Hugging Face safetensors into GGUF format, ensuring compatibility with efficient inference runtimes.
- The
gpu/convert_safetensors.pyscript handles memory-efficient loading of sharded safetensors usingsafe_open, reassembling Q, K, and V weights into a unified PyTorch state-dict. utils/convert-hf-to-gguf-bitnet.pyautomatically detects safetensors files via_is_model_safetensors()and processes them through the same quantization pipeline (supportingf32,f16,tl1, andtl2output types) used for standard PyTorch bins.- Supporting utilities in
utils/convert.pyandgpu/gguf-py/gguf.pyprovide vocabulary handling and low-level GGUF format operations.
Frequently Asked Questions
Can I convert safetensors directly to GGUF without creating an intermediate PyTorch file?
Yes. The utils/convert-hf-to-gguf-bitnet.py script automatically detects .safetensors files in your model directory using the Model._is_model_safetensors() method. When detected, it reads tensors directly using safetensors.safe_open within the get_tensors method, bypassing the need for an intermediate .pt file unless you specifically need one for debugging or manual inspection.
What quantization formats does BitNet support for GGUF output?
BitNet supports four primary output types controlled by the --outtype parameter: f32 (32-bit float), f16 (16-bit float), tl1 (ternary 1-bit), and tl2 (ternary 2-bit). The TL1 and TL2 formats are specific to BitNet's ternary quantization scheme, while f16 provides a good balance between precision and file size for standard inference.
How does BitNet handle sharded safetensors checkpoints?
The gpu/convert_safetensors.py script specifically handles sharded models through its convert_back function (lines 49-74). It constructs a TensorMap to track tensors across shards, then concatenates split weights—particularly for attention mechanisms where Q, K, and V matrices may be distributed across multiple files—into complete tensors before writing the unified checkpoint.
Why does BitNet use safetensors.safe_open instead of standard torch.load?
The conversion pipeline uses safetensors.safe_open to leverage memory-mapped file access, which allows reading tensor data without fully deserializing everything into Python objects. This approach significantly reduces RAM usage when processing large language models, prevents arbitrary code execution risks associated with pickle files, and enables efficient handling of sharded checkpoints by mapping only the required tensor segments into memory.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →