KTransformers Quantization Formats Supported: INT4, INT8, FP8, GPTQ, and IQ1_S Explained
KTransformers supports INT4, INT8, FP8 (E4M3), GPTQ-INT4, and the complete GGML IQ-series (including IQ1_S, IQ2_XXS, IQ3_XXS, and IQ4 variants), automatically selecting between AMX, AVX2, AVX-VNNI, and SYCL backends based on hardware capabilities.
The kvcache-ai/ktransformers repository implements a flexible quantization layer for large language models, particularly Mixture-of-Experts (MoE) architectures. Each KTransformers quantization format is handled through dedicated loader classes and runtime detection logic, enabling efficient inference across diverse hardware configurations.
INT4 and INT8 Quantization
KTransformers natively supports both INT4 and INT8 precision formats, including MoE-specific variants optimized for expert parallelism.
Detection Logic: The detect_quant_method() function in kt-kernel/scripts/merge_cpu_weights.py scans model directories for specific file prefixes. Lines 43-53 identify MOE_INT4_ and INT4_ prefixes for INT4 models, while lines 47-55 handle MOE_INT8_ and INT8_ prefixes for INT8 variants.
Backend Mapping: In kt-kernel/scripts/convert_cpu_weights.py (lines 695-701), these formats map to AMX backend methods through the quant_to_amx_map dictionary:
"int4"→AMXINT4"int8"→AMXINT8
# Loading INT4 weights (transparently unpacked from packed format)
from kt_kernel.python.utils.loader import SafeTensorLoader
loader = SafeTensorLoader("/path/to/int4_model")
tensor = loader.load_tensor("blk.0.ffn_up_exps.0.numa.0.weight")
print(tensor.dtype) # torch.int8 (packed INT4 storage)
FP8 (E4M3) Block-Wise Quantization
FP8 support in KTransformers implements the E4M3 format with both block-wise and per-channel scaling strategies. This format requires AVX-512 or AMX-capable hardware for optimal performance.
Implementation: The FP8SafeTensorLoader class (defined in kt-kernel/python/utils/loader.py, lines 296-306) handles FP8-specific tensor loading, reading both weight tensors and their associated scales separately.
Backend Detection: The kt-kernel/python/utils/amx.py file (lines 35-55) exposes backend variants including AVX2FP8_MOE, AMXFP8_MOE, and AMXFP8PerChannel_MOE, controlled by the _HAS_FP8_SUPPORT flag.
# Loading FP8 quantized weights
from kt_kernel.python.utils.loader import FP8SafeTensorLoader
fp8_loader = FP8SafeTensorLoader("/path/to/fp8_model")
weight = fp8_loader.load_tensor("blk.2.ffn_down_exps.1.numa.0.weight")
scale = fp8_loader.load_tensor("blk.2.ffn_down_exps.1.numa.0.scale")
print(weight.dtype, scale.dtype) # torch.uint8, torch.float32
GPTQ-INT4 for CPU and GPU
GPTQ-INT4 provides calibration-driven quantization with support for both CPU-only inference (AVX2/AVX-VNNI) and GPU offloading via SYCL.
Loader Class: When method == "GPTQ_INT4" is detected, KTransformers instantiates GPTQSafeTensorLoader (lines 900-913 in kt-kernel/python/utils/loader.py). This loader handles the asymmetric quantization data layout specific to GPTQ, including group scales and zero points.
Backend Capabilities: The amx.py module (lines 36-62) defines multiple GPTQ backends:
AVX2GPTQInt4_MOEfor baseline CPU supportAVXVNNI256GPTQInt4_MOEfor VNNI-accelerated inferenceSYCLGPTQInt4_MOEfor Intel GPU offloading
# Loading GPTQ-INT4 weights
from kt_kernel.python.utils.loader import GPTQSafeTensorLoader
gptq_loader = GPTQSafeTensorLoader("/path/to/gptq_model")
qweight = gptq_loader.load_tensor("blk.1.ffn_up_exps.0.numa.0.weight")
scales = gptq_loader.load_tensor("blk.1.ffn_up_exps.0.numa.0.scale")
print(qweight.dtype, scales.dtype) # torch.int32, torch.float32
GGML IQ-Series Quantization
KTransformers provides native support for the GGML IQ-series quantization formats, including IQ1_S, IQ2_XXS, IQ2_XS, IQ2_S, IQ3_XXS, IQ3_S, IQ4_NL, and IQ4_XS. These formats originate from the llama.cpp ecosystem and offer aggressive compression for large models.
Enum Definition: The GGMLQuantizationType enum in kt-kernel/python/utils/loader.py (lines 19-51) enumerates all supported GGML quantization types. When loading GGUF models, the GGUFReader class maps tensor type identifiers to these enum values, enabling direct IQ-series handling without AMX backend dependency.
# Loading IQ-series quantized GGUF models
from kt_kernel.python.utils.loader import GGMLQuantizationType, GGUFReader
reader = GGUFReader("/path/to/gguf_model.gguf")
tensor = reader.get_tensor("blk.0.ffn_gate_exps.0.weight")
qt = GGMLQuantizationType(tensor.quantization_type)
print(qt) # GGMLQuantizationType.IQ1_S (or IQ2_XXS, etc.)
Automatic Backend Selection
KTransformers automatically selects the optimal execution backend based on the quantization format and available hardware:
- AMX (Advanced Matrix Extensions): Used for INT4/INT8 on Intel Sapphire Rapids and newer
- AVX2/AVX-VNNI: Fallback for INT8 and GPTQ-INT4 on older CPUs
- SYCL: Enables GPTQ-INT4 offloading to Intel GPUs
- Pure Python/GGML: Handles IQ-series dequantization without specialized backends
The selection logic resides in kt-kernel/python/utils/amx.py, which sets capability flags like _HAS_FP8_SUPPORT and _HAS_AVX2_GPTQ_INT4_SUPPORT at module initialization.
Summary
- INT4/INT8: Native support via
SafeTensorLoaderwith AMX acceleration and MoE-specific variants detected by filename prefixes inmerge_cpu_weights.py - FP8: E4M3 format requiring AVX-512 or AMX hardware, handled by
FP8SafeTensorLoaderwith separate scale tensors - GPTQ-INT4: Calibration-based quantization supporting AVX2, AVX-VNNI, and SYCL GPU backends via
GPTQSafeTensorLoader - IQ-Series: Full GGML compatibility including IQ1_S, IQ2, IQ3, and IQ4 variants through
GGMLQuantizationTypeenum andGGUFReader - Hardware Optimization: Automatic backend selection between AMX, AVX2, AVX-VNNI, and SYCL based on capability detection in
amx.py
Frequently Asked Questions
What is the difference between standard INT4 and GPTQ-INT4 in KTransformers?
Standard INT4 uses simple symmetric quantization suitable for MoE expert layers, while GPTQ-INT4 employs one-shot weight quantization with calibration data, storing asymmetric scales and zero-points. GPTQ-INT4 uses GPTQSafeTensorLoader instead of the base SafeTensorLoader and supports GPU offloading via SYCL, whereas standard INT4 targets CPU AMX acceleration.
Does KTransformers support IQ1_S quantization for extreme compression?
Yes, KTransformers supports IQ1_S and the complete GGML IQ-series through the GGMLQuantizationType enum in kt-kernel/python/utils/loader.py (lines 19-51). When loading GGUF files with IQ1_S tensors, the GGUFReader class handles dequantization natively without requiring AMX or other specialized backends.
How does KTransformers automatically select the quantization backend?
Backend selection occurs in kt-kernel/python/utils/amx.py, which probes CPU features at import time to set flags like _HAS_FP8_SUPPORT, _HAS_AVX2_GPTQ_INT4_SUPPORT, and _HAS_AMX_INT8_SUPPORT. The loader classes then dispatch to the appropriate implementation (AMXINT4, AVX2FP8_MOE, etc.) based on these capability flags combined with the detected quantization method.
Can I mix different quantization formats in the same KTransformers model?
While KTransformers supports loading different formats through their respective loader classes (SafeTensorLoader, FP8SafeTensorLoader, GPTQSafeTensorLoader, GGUFReader), a single model checkpoint typically uses one quantization scheme throughout. The convert_cpu_weights.py script handles conversion between formats, but runtime mixing requires manual tensor management.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →