# How to Convert Hugging Face Safetensors to GGUF for BitNet: Complete Guide

> Learn how to convert Hugging Face Safetensors to GGUF for BitNet models. Follow our guide to reassemble sharded weights and create quantized GGUF files for efficient AI.

- Repository: [Microsoft/BitNet](https://github.com/microsoft/BitNet)
- Tags: how-to-guide
- Published: 2026-03-13

---

**BitNet provides a two-stage pipeline that converts Hugging Face safetensors checkpoints to GGUF format using [`gpu/convert_safetensors.py`](https://github.com/microsoft/BitNet/blob/main/gpu/convert_safetensors.py) to reassemble sharded weights into PyTorch format, followed by [`utils/convert-hf-to-gguf-bitnet.py`](https://github.com/microsoft/BitNet/blob/main/utils/convert-hf-to-gguf-bitnet.py) to generate the final quantized GGUF file.**

The microsoft/BitNet repository includes specialized utilities designed to bridge the gap between Hugging Face's safetensors format and the GGUF format required for efficient BitNet inference. This guide explains the complete workflow to convert Hugging Face safetensors to GGUF for BitNet, covering both the automatic detection capabilities and the optional manual reassembly process.

## Understanding the Two-Stage Conversion Pipeline

BitNet's conversion process is architected as a two-stage pipeline that handles the memory-mapped nature of safetensors while maintaining compatibility with the standard GGUF generation workflow.

### Stage 1: Reassembling Safetensors into PyTorch Format

The first stage addresses the safetensors format's memory-mapped, read-only structure. In [`gpu/convert_safetensors.py`](https://github.com/microsoft/BitNet/blob/main/gpu/convert_safetensors.py), the `convert_back` function (lines 49-74) uses `safetensors.safe_open` to read tensors without fully deserializing them into Python objects, significantly reducing RAM usage during conversion.

This script specifically handles sharded checkpoints by building a TensorMap to track which tensors belong to which shards. It then concatenates split weights—particularly for Q, K, and V projection matrices—into complete tensors before writing a unified PyTorch state-dict (`.pt`) that matches the layout expected by the GGUF converter.

### Stage 2: Converting to GGUF Format

The second stage utilizes [`utils/convert-hf-to-gguf-bitnet.py`](https://github.com/microsoft/BitNet/blob/main/utils/convert-hf-to-gguf-bitnet.py), which serves as the core Hugging Face to GGUF conversion engine. This script automatically detects whether the model directory contains `.safetensors` files via the `Model._is_model_safetensors()` method.

When safetensors are detected, the `get_tensors` method (lines 58-90) switches to using `safetensors.safe_open` instead of `torch.load`. The remainder of the conversion process—including tensor renaming, optional TL1/TL2 quantization, and vocabulary handling—proceeds identically to the standard `.bin` workflow, ensuring consistent GGUF output regardless of the original checkpoint format.

## Step-by-Step Conversion Workflow

### 1. Install BitNet Conversion Utilities

Begin by cloning the repository and installing the required Python packages:

```bash
git clone https://github.com/microsoft/BitNet.git
cd BitNet
pip install -r requirements.txt

```

The requirements include `safetensors`, `torch`, `gguf-py`, `numpy`, and `einops`.

### 2. (Optional) Convert Safetensors to PyTorch Checkpoint

If you require an intermediate PyTorch state-dict for inspection or debugging, use the dedicated safetensors converter:

```bash
python gpu/convert_safetensors.py \
  --safetensors_file path/to/model.safetensors \
  --output model.pt \
  --model_name 2B

```

Select a `--model_name` that matches your model configuration (e.g., `2B`). This produces a `model.pt` file containing the reassembled weights.

### 3. Run the GGUF Converter

Execute the main conversion script, pointing it to your model directory containing the safetensors files:

```bash
python utils/convert-hf-to-gguf-bitnet.py \
  --model path/to/model_dir \
  --outtype f16 \
  --outfile model.gguf

```

Available `--outtype` options include `f32` (full precision), `f16` (half precision), `tl1` (ternary 1-bit), and `tl2` (ternary 2-bit). The script automatically detects `.safetensors` files and processes them through the same pipeline as PyTorch bins.

### 4. Verify the Output

Confirm the conversion succeeded by inspecting the GGUF metadata:

```bash
gguf-tool info model.gguf

```

Alternatively, verify programmatically:

```python
import gguf

gguf_file = gguf.GGUFReader('model.gguf')
print('Arch:', gguf_file.metadata['model_arch'])
print('Layers:', gguf_file.metadata['block_count'])
print('Quantisation:', gguf_file.metadata['file_type'])

```

## Code Examples

### Direct Conversion from a Safetensors Directory

```bash

# Assume the HF repo was cloned to ./my-bitnet-model

# The model directory contains model.safetensors + config.json etc.

python utils/convert-hf-to-gguf-bitnet.py \
    --model ./my-bitnet-model \
    --outtype tl2 \
    --outfile ./my-bitnet-model/bitnet.gguf

```

### Using an Intermediate PyTorch Checkpoint

```bash

# Step 1: safetensors → torch checkpoint

python gpu/convert_safetensors.py \
    --safetensors_file ./my-bitnet-model/model.safetensors \
    --output ./my-bitnet-model/model.pt \
    --model_name 2B

# Step 2: torch checkpoint → GGUF

python utils/convert-hf-to-gguf-bitnet.py \
    --model ./my-bitnet-model \
    --outtype f16 \
    --outfile ./my-bitnet-model/bitnet-f16.gguf

```

### Inspecting the Generated GGUF File

```python
import gguf

gguf_file = gguf.GGUFReader('bitnet.gguf')
print('Architecture:', gguf_file.metadata['model_arch'])
print('Block count:', gguf_file.metadata['block_count'])
print('File type:', gguf_file.metadata['file_type'])

```

## Key Implementation Files

The conversion pipeline relies on these specific files in the microsoft/BitNet repository:

| File | Role | Key Components |
|------|------|----------------|
| [`gpu/convert_safetensors.py`](https://github.com/microsoft/BitNet/blob/main/gpu/convert_safetensors.py) | Reads `.safetensors`, reassembles sharded model parts, writes PyTorch checkpoint. | `convert_back` function (lines 49-74), TensorMap construction |
| [`utils/convert-hf-to-gguf-bitnet.py`](https://github.com/microsoft/BitNet/blob/main/utils/convert-hf-to-gguf-bitnet.py) | Core HF-to-GGUF conversion engine; auto-detects safetensors and handles quantization. | `Model._is_model_safetensors()`, `get_tensors` method (lines 58-90) |
| [`utils/convert.py`](https://github.com/microsoft/BitNet/blob/main/utils/convert.py) | Helper utilities shared across model families for vocabulary handling and tensor permutations. | `LlamaHfVocab`, `permute` functions |
| [`gpu/gguf-py/gguf.py`](https://github.com/microsoft/BitNet/blob/main/gpu/gguf-py/gguf.py) | Vendored GGUF reader/writer implementation used throughout the conversion process. | `GGUFReader`, `GGUFWriter` classes |
| [`gpu/convert_checkpoint.py`](https://github.com/microsoft/BitNet/blob/main/gpu/convert_checkpoint.py) | Alternative checkpoint conversion entry point using the same Model abstraction. | Checkpoint loading and validation logic |

## Summary

- BitNet implements a **two-stage conversion pipeline** to transform Hugging Face safetensors into GGUF format, ensuring compatibility with efficient inference runtimes.
- The **[`gpu/convert_safetensors.py`](https://github.com/microsoft/BitNet/blob/main/gpu/convert_safetensors.py)** script handles memory-efficient loading of sharded safetensors using `safe_open`, reassembling Q, K, and V weights into a unified PyTorch state-dict.
- **[`utils/convert-hf-to-gguf-bitnet.py`](https://github.com/microsoft/BitNet/blob/main/utils/convert-hf-to-gguf-bitnet.py)** automatically detects safetensors files via `_is_model_safetensors()` and processes them through the same quantization pipeline (supporting `f32`, `f16`, `tl1`, and `tl2` output types) used for standard PyTorch bins.
- Supporting utilities in [`utils/convert.py`](https://github.com/microsoft/BitNet/blob/main/utils/convert.py) and [`gpu/gguf-py/gguf.py`](https://github.com/microsoft/BitNet/blob/main/gpu/gguf-py/gguf.py) provide vocabulary handling and low-level GGUF format operations.

## Frequently Asked Questions

### Can I convert safetensors directly to GGUF without creating an intermediate PyTorch file?

Yes. The [`utils/convert-hf-to-gguf-bitnet.py`](https://github.com/microsoft/BitNet/blob/main/utils/convert-hf-to-gguf-bitnet.py) script automatically detects `.safetensors` files in your model directory using the `Model._is_model_safetensors()` method. When detected, it reads tensors directly using `safetensors.safe_open` within the `get_tensors` method, bypassing the need for an intermediate `.pt` file unless you specifically need one for debugging or manual inspection.

### What quantization formats does BitNet support for GGUF output?

BitNet supports four primary output types controlled by the `--outtype` parameter: `f32` (32-bit float), `f16` (16-bit float), `tl1` (ternary 1-bit), and `tl2` (ternary 2-bit). The TL1 and TL2 formats are specific to BitNet's ternary quantization scheme, while f16 provides a good balance between precision and file size for standard inference.

### How does BitNet handle sharded safetensors checkpoints?

The [`gpu/convert_safetensors.py`](https://github.com/microsoft/BitNet/blob/main/gpu/convert_safetensors.py) script specifically handles sharded models through its `convert_back` function (lines 49-74). It constructs a TensorMap to track tensors across shards, then concatenates split weights—particularly for attention mechanisms where Q, K, and V matrices may be distributed across multiple files—into complete tensors before writing the unified checkpoint.

### Why does BitNet use safetensors.safe_open instead of standard torch.load?

The conversion pipeline uses `safetensors.safe_open` to leverage memory-mapped file access, which allows reading tensor data without fully deserializing everything into Python objects. This approach significantly reduces RAM usage when processing large language models, prevents arbitrary code execution risks associated with pickle files, and enables efficient handling of sharded checkpoints by mapping only the required tensor segments into memory.