# How to Convert and Optimize Model Weights for KTransformers Inference Using convert_cpu_weights.py

> Learn how to convert and optimize model weights for KTransformers CPU inference using convert_cpu_weights.py. Quantize weights to INT4 or INT8 for faster performance.

- Repository: [kvcache.ai/ktransformers](https://github.com/kvcache-ai/ktransformers)
- Tags: how-to-guide
- Published: 2026-07-26

---

**The [`convert_cpu_weights.py`](https://github.com/kvcache-ai/ktransformers/blob/main/convert_cpu_weights.py) script prepares FP16/BF16 model checkpoints for high-throughput CPU inference by quantizing weights to INT4 or INT8 and restructuring tensors into the KTransformers runtime format.**

The kvcache-ai/ktransformers repository provides specialized CPU inference engines for large language models that require specifically formatted weight tensors to maximize memory bandwidth efficiency. Converting full-precision checkpoints using the official conversion utilities is a mandatory preprocessing step that reduces model size by 50-75% while preserving generation quality through optimized quantization kernels.

## Prerequisites and Installation

Before running the conversion utility, ensure your environment meets the following requirements:

- **KTransformers kernel libraries**: Install the `kt-kernel` package by running `pip install -e .` from the repository root.
- **PyTorch with CPU support**: The conversion script relies on `torch.load` and HuggingFace `safetensors` for tensor I/O operations.
- **Source checkpoint**: A directory containing the original model weights in BF16, FP8, or FP16 format (either `.bin` PyTorch files or `.safetensors` files).

Verify the installation by checking that [`kt-kernel/scripts/convert_cpu_weights.py`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/scripts/convert_cpu_weights.py) exists in your local clone of the repository.

## Understanding the Conversion Pipeline

The conversion process implemented in [`kt-kernel/scripts/convert_cpu_weights.py`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/scripts/convert_cpu_weights.py) performs three critical operations:

1. **Tensor Loading**: Reads original weights using `torch.load` for PyTorch binaries or `safetensors.load_file` for Safetensors formats.
2. **Quantization**: Applies INT4 or INT8 quantization routines that map floating-point values to discrete integers while minimizing perplexity degradation.
3. **Layout Optimization**: Rewrites tensors into the blocked format expected by the CPU inference backend in [`ktransformers/ktransformers_ext/triton/fp8gemm.py`](https://github.com/kvcache-ai/ktransformers/blob/main/ktransformers/ktransformers_ext/triton/fp8gemm.py), which enables efficient vectorized operations on x86 AVX2/AVX-512 hardware.

The script preserves all metadata files—including [`config.json`](https://github.com/kvcache-ai/ktransformers/blob/main/config.json) and tokenizer configurations—ensuring the output directory remains compatible with standard HuggingFace `from_pretrained` workflows.

## Step-by-Step Weight Conversion

### Prepare the Source Checkpoint

Organize your original model files in a dedicated directory. For example, if converting Qwen2-7B:

```bash
mkdir -p /data/models/Qwen2-7B-original

# Copy model.safetensors, config.json, tokenizer.json here

```

Ensure the directory contains the appropriate weight files and that you have write permissions for the destination path.

### Execute the Conversion Command

Invoke the script with the following required arguments:

```bash
python kt-kernel/scripts/convert_cpu_weights.py \
    --input-path /data/models/Qwen2-7B-original/ \
    --input-type bf16 \
    --output /data/models/Qwen2-7B-INT4/ \
    --quant-method int4 \
    --cpuinfer-threads 60 \
    --threadpool-count 2 \
    --overwrite

```

**Parameter breakdown:**
- **`--input-path`**: Absolute path to the source checkpoint directory.
- **`--input-type`**: Original precision (`bf16`, `fp8`, or `fp16`).
- **`--output`**: Target directory for the quantized checkpoint.
- **`--quant-method`**: Target quantization (`int4` for maximum compression, `int8` for higher fidelity).
- **`--overwrite`**: Forces replacement of existing files in the output directory.

### Configure Thread Pool Settings

The **`--cpuinfer-threads`** and **`--threadpool-count`** parameters control parallelism during inference, not conversion. However, setting them during conversion embeds optimal configuration metadata into the checkpoint:

- **`--cpuinfer-threads`**: Total CPU cores allocated to the inference engine (typically set to physical core count).
- **`--threadpool-count`**: Number of asynchronous generation threads; values of 2-4 optimize throughput for batch processing.

For a 64-core server, `--cpuinfer-threads 60 --threadpool-count 2` reserves four cores for system operations while saturating the remainder with compute threads.

## Advanced Quantization Options

### Handling MoE Architectures

For Mixture-of-Experts (MoE) models containing sparse expert layers, first convert the checkpoint to BF16 precision using [`kt-kernel/scripts/convert_moe_to_bf16.py`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/scripts/convert_moe_to_bf16.py) before applying INT4 quantization:

```bash
python kt-kernel/scripts/convert_moe_to_bf16.py \
    --input-path /data/models/Mixtral-MoE/ \
    --output /data/models/Mixtral-BF16/

python kt-kernel/scripts/convert_cpu_weights.py \
    --input-path /data/models/Mixtral-BF16/ \
    --input-type bf16 \
    --quant-method int4 \
    --output /data/models/Mixtral-INT4/

```

This two-stage process ensures expert routing weights maintain sufficient precision before aggressive quantization.

## Verification and Debugging

After conversion, validate tensor integrity using the comparison utility:

```bash
python kt-kernel/scripts/compare_weights.py \
    --original /data/models/Qwen2-7B-original/ \
    --converted /data/models/Qwen2-7B-INT4/

```

This script checks that:
- All layer names exist in both checkpoints.
- Quantized tensors match expected shapes.
- No NaN or Inf values corrupted the conversion process.

If discrepancies appear, verify that `--input-type` matches the actual precision of the source files (mismatches cause loading errors in `torch.load`).

## Loading Converted Weights for Inference

The output directory produced by [`convert_cpu_weights.py`](https://github.com/kvcache-ai/ktransformers/blob/main/convert_cpu_weights.py) is immediately consumable by the KTransformers runtime. Load the optimized model in Python:

```python
from ktransformers import KTransformer

model = KTransformer.from_pretrained(
    "/data/models/Qwen2-7B-INT4/",
    device="cpu"
)

output = model.generate(
    "Explain quantum computing concepts:",
    max_new_tokens=512,
    temperature=0.7
)
print(output)

```

The `KTransformer` class automatically detects the INT4/INT8 format and dispatches to optimized kernels in [`ktransformers/ktransformers_ext/triton/fp8gemm.py`](https://github.com/kvcache-ai/ktransformers/blob/main/ktransformers/ktransformers_ext/triton/fp8gemm.py), achieving inference speeds comparable to GPU-based quantization on high-core-count CPUs.

## Summary

- **[`convert_cpu_weights.py`](https://github.com/kvcache-ai/ktransformers/blob/main/convert_cpu_weights.py)** transforms standard checkpoints into KTransformers-compatible INT4/INT8 formats, located at [`kt-kernel/scripts/convert_cpu_weights.py`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/scripts/convert_cpu_weights.py) in the repository.
- The script supports both PyTorch `.bin` and Safetensors formats via `torch.load` and `safetensors.load_file` APIs.
- Key arguments include `--quant-method` (int4/int8), `--cpuinfer-threads`, and `--threadpool-count` for performance tuning.
- MoE models require pre-processing with [`convert_moe_to_bf16.py`](https://github.com/kvcache-ai/ktransformers/blob/main/convert_moe_to_bf16.py) before quantization.
- Verify conversions using [`kt-kernel/scripts/compare_weights.py`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/scripts/compare_weights.py) before deployment.
- Converted checkpoints load directly into `KTransformer` objects without additional preprocessing.

## Frequently Asked Questions

### What quantization formats does convert_cpu_weights.py support?

The script supports **INT4** and **INT8** quantization methods. INT4 provides maximum compression (approximately 4 bits per weight) for memory-constrained environments, while INT8 (8 bits per weight) offers higher accuracy for critical applications. Both formats utilize optimized GEMM kernels in the KTransformers backend.

### How do I determine optimal values for --cpuinfer-threads?

Set `--cpuinfer-threads` to the number of physical CPU cores minus two to four cores reserved for the operating system and Python overhead. For hyper-threaded systems, use the physical core count rather than logical threads to avoid context-switching penalties. Values between 32 and 64 typically maximize throughput on server-grade processors.

### Can I convert models stored as Safetensors?

Yes. The conversion script automatically detects file extensions and uses `safetensors.load_file` for `.safetensors` files or `torch.load` for `.bin` files. Ensure the `--input-path` directory contains the appropriate format—the script processes both types without requiring format-specific flags.

### Why are my converted weights failing to load in KTransformers?

Loading failures usually indicate a mismatch between the declared `--input-type` and the actual source precision, or missing metadata files in the output directory. Verify that [`config.json`](https://github.com/kvcache-ai/ktransformers/blob/main/config.json) was copied to the output folder, and run [`compare_weights.py`](https://github.com/kvcache-ai/ktransformers/blob/main/compare_weights.py) to check for tensor corruption. If using INT4, ensure your KTransformers build includes the [`fp8gemm.py`](https://github.com/kvcache-ai/ktransformers/blob/main/fp8gemm.py) backend components required for 4-bit inference.