How to Convert and Optimize Model Weights for KTransformers Inference Using convert_cpu_weights.py
The convert_cpu_weights.py script prepares FP16/BF16 model checkpoints for high-throughput CPU inference by quantizing weights to INT4 or INT8 and restructuring tensors into the KTransformers runtime format.
The kvcache-ai/ktransformers repository provides specialized CPU inference engines for large language models that require specifically formatted weight tensors to maximize memory bandwidth efficiency. Converting full-precision checkpoints using the official conversion utilities is a mandatory preprocessing step that reduces model size by 50-75% while preserving generation quality through optimized quantization kernels.
Prerequisites and Installation
Before running the conversion utility, ensure your environment meets the following requirements:
- KTransformers kernel libraries: Install the
kt-kernelpackage by runningpip install -e .from the repository root. - PyTorch with CPU support: The conversion script relies on
torch.loadand HuggingFacesafetensorsfor tensor I/O operations. - Source checkpoint: A directory containing the original model weights in BF16, FP8, or FP16 format (either
.binPyTorch files or.safetensorsfiles).
Verify the installation by checking that kt-kernel/scripts/convert_cpu_weights.py exists in your local clone of the repository.
Understanding the Conversion Pipeline
The conversion process implemented in kt-kernel/scripts/convert_cpu_weights.py performs three critical operations:
- Tensor Loading: Reads original weights using
torch.loadfor PyTorch binaries orsafetensors.load_filefor Safetensors formats. - Quantization: Applies INT4 or INT8 quantization routines that map floating-point values to discrete integers while minimizing perplexity degradation.
- Layout Optimization: Rewrites tensors into the blocked format expected by the CPU inference backend in
ktransformers/ktransformers_ext/triton/fp8gemm.py, which enables efficient vectorized operations on x86 AVX2/AVX-512 hardware.
The script preserves all metadata files—including config.json and tokenizer configurations—ensuring the output directory remains compatible with standard HuggingFace from_pretrained workflows.
Step-by-Step Weight Conversion
Prepare the Source Checkpoint
Organize your original model files in a dedicated directory. For example, if converting Qwen2-7B:
mkdir -p /data/models/Qwen2-7B-original
# Copy model.safetensors, config.json, tokenizer.json here
Ensure the directory contains the appropriate weight files and that you have write permissions for the destination path.
Execute the Conversion Command
Invoke the script with the following required arguments:
python kt-kernel/scripts/convert_cpu_weights.py \
--input-path /data/models/Qwen2-7B-original/ \
--input-type bf16 \
--output /data/models/Qwen2-7B-INT4/ \
--quant-method int4 \
--cpuinfer-threads 60 \
--threadpool-count 2 \
--overwrite
Parameter breakdown:
--input-path: Absolute path to the source checkpoint directory.--input-type: Original precision (bf16,fp8, orfp16).--output: Target directory for the quantized checkpoint.--quant-method: Target quantization (int4for maximum compression,int8for higher fidelity).--overwrite: Forces replacement of existing files in the output directory.
Configure Thread Pool Settings
The --cpuinfer-threads and --threadpool-count parameters control parallelism during inference, not conversion. However, setting them during conversion embeds optimal configuration metadata into the checkpoint:
--cpuinfer-threads: Total CPU cores allocated to the inference engine (typically set to physical core count).--threadpool-count: Number of asynchronous generation threads; values of 2-4 optimize throughput for batch processing.
For a 64-core server, --cpuinfer-threads 60 --threadpool-count 2 reserves four cores for system operations while saturating the remainder with compute threads.
Advanced Quantization Options
Handling MoE Architectures
For Mixture-of-Experts (MoE) models containing sparse expert layers, first convert the checkpoint to BF16 precision using kt-kernel/scripts/convert_moe_to_bf16.py before applying INT4 quantization:
python kt-kernel/scripts/convert_moe_to_bf16.py \
--input-path /data/models/Mixtral-MoE/ \
--output /data/models/Mixtral-BF16/
python kt-kernel/scripts/convert_cpu_weights.py \
--input-path /data/models/Mixtral-BF16/ \
--input-type bf16 \
--quant-method int4 \
--output /data/models/Mixtral-INT4/
This two-stage process ensures expert routing weights maintain sufficient precision before aggressive quantization.
Verification and Debugging
After conversion, validate tensor integrity using the comparison utility:
python kt-kernel/scripts/compare_weights.py \
--original /data/models/Qwen2-7B-original/ \
--converted /data/models/Qwen2-7B-INT4/
This script checks that:
- All layer names exist in both checkpoints.
- Quantized tensors match expected shapes.
- No NaN or Inf values corrupted the conversion process.
If discrepancies appear, verify that --input-type matches the actual precision of the source files (mismatches cause loading errors in torch.load).
Loading Converted Weights for Inference
The output directory produced by convert_cpu_weights.py is immediately consumable by the KTransformers runtime. Load the optimized model in Python:
from ktransformers import KTransformer
model = KTransformer.from_pretrained(
"/data/models/Qwen2-7B-INT4/",
device="cpu"
)
output = model.generate(
"Explain quantum computing concepts:",
max_new_tokens=512,
temperature=0.7
)
print(output)
The KTransformer class automatically detects the INT4/INT8 format and dispatches to optimized kernels in ktransformers/ktransformers_ext/triton/fp8gemm.py, achieving inference speeds comparable to GPU-based quantization on high-core-count CPUs.
Summary
convert_cpu_weights.pytransforms standard checkpoints into KTransformers-compatible INT4/INT8 formats, located atkt-kernel/scripts/convert_cpu_weights.pyin the repository.- The script supports both PyTorch
.binand Safetensors formats viatorch.loadandsafetensors.load_fileAPIs. - Key arguments include
--quant-method(int4/int8),--cpuinfer-threads, and--threadpool-countfor performance tuning. - MoE models require pre-processing with
convert_moe_to_bf16.pybefore quantization. - Verify conversions using
kt-kernel/scripts/compare_weights.pybefore deployment. - Converted checkpoints load directly into
KTransformerobjects without additional preprocessing.
Frequently Asked Questions
What quantization formats does convert_cpu_weights.py support?
The script supports INT4 and INT8 quantization methods. INT4 provides maximum compression (approximately 4 bits per weight) for memory-constrained environments, while INT8 (8 bits per weight) offers higher accuracy for critical applications. Both formats utilize optimized GEMM kernels in the KTransformers backend.
How do I determine optimal values for --cpuinfer-threads?
Set --cpuinfer-threads to the number of physical CPU cores minus two to four cores reserved for the operating system and Python overhead. For hyper-threaded systems, use the physical core count rather than logical threads to avoid context-switching penalties. Values between 32 and 64 typically maximize throughput on server-grade processors.
Can I convert models stored as Safetensors?
Yes. The conversion script automatically detects file extensions and uses safetensors.load_file for .safetensors files or torch.load for .bin files. Ensure the --input-path directory contains the appropriate format—the script processes both types without requiring format-specific flags.
Why are my converted weights failing to load in KTransformers?
Loading failures usually indicate a mismatch between the declared --input-type and the actual source precision, or missing metadata files in the output directory. Verify that config.json was copied to the output folder, and run compare_weights.py to check for tensor corruption. If using INT4, ensure your KTransformers build includes the fp8gemm.py backend components required for 4-bit inference.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →