# KTransformers Quantization Formats Supported: INT4, INT8, FP8, GPTQ, and IQ1_S Explained

> Explore KTransformers quantization formats like INT4, INT8, FP8, GPTQ, and IQ1_S. Discover automatic hardware backend selection for optimal performance.

- Repository: [kvcache.ai/ktransformers](https://github.com/kvcache-ai/ktransformers)
- Tags: deep-dive
- Published: 2026-07-26

---

**KTransformers supports INT4, INT8, FP8 (E4M3), GPTQ-INT4, and the complete GGML IQ-series (including IQ1_S, IQ2_XXS, IQ3_XXS, and IQ4 variants), automatically selecting between AMX, AVX2, AVX-VNNI, and SYCL backends based on hardware capabilities.**

The kvcache-ai/ktransformers repository implements a flexible quantization layer for large language models, particularly Mixture-of-Experts (MoE) architectures. Each KTransformers quantization format is handled through dedicated loader classes and runtime detection logic, enabling efficient inference across diverse hardware configurations.

## INT4 and INT8 Quantization

KTransformers natively supports both **INT4** and **INT8** precision formats, including MoE-specific variants optimized for expert parallelism.

**Detection Logic:** The `detect_quant_method()` function in [`kt-kernel/scripts/merge_cpu_weights.py`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/scripts/merge_cpu_weights.py) scans model directories for specific file prefixes. Lines 43-53 identify `MOE_INT4_` and `INT4_` prefixes for INT4 models, while lines 47-55 handle `MOE_INT8_` and `INT8_` prefixes for INT8 variants.

**Backend Mapping:** In [`kt-kernel/scripts/convert_cpu_weights.py`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/scripts/convert_cpu_weights.py) (lines 695-701), these formats map to AMX backend methods through the `quant_to_amx_map` dictionary:
- `"int4"` → `AMXINT4`
- `"int8"` → `AMXINT8`

```python

# Loading INT4 weights (transparently unpacked from packed format)

from kt_kernel.python.utils.loader import SafeTensorLoader

loader = SafeTensorLoader("/path/to/int4_model")
tensor = loader.load_tensor("blk.0.ffn_up_exps.0.numa.0.weight")
print(tensor.dtype)  # torch.int8 (packed INT4 storage)

```

## FP8 (E4M3) Block-Wise Quantization

**FP8** support in KTransformers implements the E4M3 format with both block-wise and per-channel scaling strategies. This format requires AVX-512 or AMX-capable hardware for optimal performance.

**Implementation:** The `FP8SafeTensorLoader` class (defined in [`kt-kernel/python/utils/loader.py`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/python/utils/loader.py), lines 296-306) handles FP8-specific tensor loading, reading both weight tensors and their associated scales separately.

**Backend Detection:** The [`kt-kernel/python/utils/amx.py`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/python/utils/amx.py) file (lines 35-55) exposes backend variants including `AVX2FP8_MOE`, `AMXFP8_MOE`, and `AMXFP8PerChannel_MOE`, controlled by the `_HAS_FP8_SUPPORT` flag.

```python

# Loading FP8 quantized weights

from kt_kernel.python.utils.loader import FP8SafeTensorLoader

fp8_loader = FP8SafeTensorLoader("/path/to/fp8_model")
weight = fp8_loader.load_tensor("blk.2.ffn_down_exps.1.numa.0.weight")
scale = fp8_loader.load_tensor("blk.2.ffn_down_exps.1.numa.0.scale")
print(weight.dtype, scale.dtype)  # torch.uint8, torch.float32

```

## GPTQ-INT4 for CPU and GPU

**GPTQ-INT4** provides calibration-driven quantization with support for both CPU-only inference (AVX2/AVX-VNNI) and GPU offloading via SYCL.

**Loader Class:** When `method == "GPTQ_INT4"` is detected, KTransformers instantiates `GPTQSafeTensorLoader` (lines 900-913 in [`kt-kernel/python/utils/loader.py`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/python/utils/loader.py)). This loader handles the asymmetric quantization data layout specific to GPTQ, including group scales and zero points.

**Backend Capabilities:** The [`amx.py`](https://github.com/kvcache-ai/ktransformers/blob/main/amx.py) module (lines 36-62) defines multiple GPTQ backends:
- `AVX2GPTQInt4_MOE` for baseline CPU support
- `AVXVNNI256GPTQInt4_MOE` for VNNI-accelerated inference
- `SYCLGPTQInt4_MOE` for Intel GPU offloading

```python

# Loading GPTQ-INT4 weights

from kt_kernel.python.utils.loader import GPTQSafeTensorLoader

gptq_loader = GPTQSafeTensorLoader("/path/to/gptq_model")
qweight = gptq_loader.load_tensor("blk.1.ffn_up_exps.0.numa.0.weight")
scales = gptq_loader.load_tensor("blk.1.ffn_up_exps.0.numa.0.scale")
print(qweight.dtype, scales.dtype)  # torch.int32, torch.float32

```

## GGML IQ-Series Quantization

KTransformers provides native support for the **GGML IQ-series** quantization formats, including **IQ1_S**, IQ2_XXS, IQ2_XS, IQ2_S, IQ3_XXS, IQ3_S, IQ4_NL, and IQ4_XS. These formats originate from the llama.cpp ecosystem and offer aggressive compression for large models.

**Enum Definition:** The `GGMLQuantizationType` enum in [`kt-kernel/python/utils/loader.py`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/python/utils/loader.py) (lines 19-51) enumerates all supported GGML quantization types. When loading GGUF models, the `GGUFReader` class maps tensor type identifiers to these enum values, enabling direct IQ-series handling without AMX backend dependency.

```python

# Loading IQ-series quantized GGUF models

from kt_kernel.python.utils.loader import GGMLQuantizationType, GGUFReader

reader = GGUFReader("/path/to/gguf_model.gguf")
tensor = reader.get_tensor("blk.0.ffn_gate_exps.0.weight")
qt = GGMLQuantizationType(tensor.quantization_type)
print(qt)  # GGMLQuantizationType.IQ1_S (or IQ2_XXS, etc.)

```

## Automatic Backend Selection

KTransformers automatically selects the optimal execution backend based on the quantization format and available hardware:

- **AMX (Advanced Matrix Extensions):** Used for INT4/INT8 on Intel Sapphire Rapids and newer
- **AVX2/AVX-VNNI:** Fallback for INT8 and GPTQ-INT4 on older CPUs
- **SYCL:** Enables GPTQ-INT4 offloading to Intel GPUs
- **Pure Python/GGML:** Handles IQ-series dequantization without specialized backends

The selection logic resides in [`kt-kernel/python/utils/amx.py`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/python/utils/amx.py), which sets capability flags like `_HAS_FP8_SUPPORT` and `_HAS_AVX2_GPTQ_INT4_SUPPORT` at module initialization.

## Summary

- **INT4/INT8:** Native support via `SafeTensorLoader` with AMX acceleration and MoE-specific variants detected by filename prefixes in [`merge_cpu_weights.py`](https://github.com/kvcache-ai/ktransformers/blob/main/merge_cpu_weights.py)
- **FP8:** E4M3 format requiring AVX-512 or AMX hardware, handled by `FP8SafeTensorLoader` with separate scale tensors
- **GPTQ-INT4:** Calibration-based quantization supporting AVX2, AVX-VNNI, and SYCL GPU backends via `GPTQSafeTensorLoader`
- **IQ-Series:** Full GGML compatibility including IQ1_S, IQ2, IQ3, and IQ4 variants through `GGMLQuantizationType` enum and `GGUFReader`
- **Hardware Optimization:** Automatic backend selection between AMX, AVX2, AVX-VNNI, and SYCL based on capability detection in [`amx.py`](https://github.com/kvcache-ai/ktransformers/blob/main/amx.py)

## Frequently Asked Questions

### What is the difference between standard INT4 and GPTQ-INT4 in KTransformers?

Standard INT4 uses simple symmetric quantization suitable for MoE expert layers, while GPTQ-INT4 employs one-shot weight quantization with calibration data, storing asymmetric scales and zero-points. GPTQ-INT4 uses `GPTQSafeTensorLoader` instead of the base `SafeTensorLoader` and supports GPU offloading via SYCL, whereas standard INT4 targets CPU AMX acceleration.

### Does KTransformers support IQ1_S quantization for extreme compression?

Yes, KTransformers supports **IQ1_S** and the complete GGML IQ-series through the `GGMLQuantizationType` enum in [`kt-kernel/python/utils/loader.py`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/python/utils/loader.py) (lines 19-51). When loading GGUF files with IQ1_S tensors, the `GGUFReader` class handles dequantization natively without requiring AMX or other specialized backends.

### How does KTransformers automatically select the quantization backend?

Backend selection occurs in [`kt-kernel/python/utils/amx.py`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/python/utils/amx.py), which probes CPU features at import time to set flags like `_HAS_FP8_SUPPORT`, `_HAS_AVX2_GPTQ_INT4_SUPPORT`, and `_HAS_AMX_INT8_SUPPORT`. The loader classes then dispatch to the appropriate implementation (AMXINT4, AVX2FP8_MOE, etc.) based on these capability flags combined with the detected quantization method.

### Can I mix different quantization formats in the same KTransformers model?

While KTransformers supports loading different formats through their respective loader classes (`SafeTensorLoader`, `FP8SafeTensorLoader`, `GPTQSafeTensorLoader`, `GGUFReader`), a single model checkpoint typically uses one quantization scheme throughout. The [`convert_cpu_weights.py`](https://github.com/kvcache-ai/ktransformers/blob/main/convert_cpu_weights.py) script handles conversion between formats, but runtime mixing requires manual tensor management.