KTransformers Supported Models: DeepSeek, Qwen, GLM-4, MiniMax & Kimi-K2
KTransformers natively supports inference for DeepSeek, Qwen (2/3/3.5), GLM-4, MiniMax, and Kimi-K2 MoE architectures through a centralized model registry and automatic architecture detection system.
KTransformers is an open-source inference engine specifically optimized for Mixture-of-Experts (MoE) large language models. The framework maintains a comprehensive model registry at kt-kernel/python/cli/utils/model_registry.py alongside architecture detection logic in kt-kernel/python/sft/arch.py that automatically maps weight layouts to optimized kernels. According to the kvcache-ai/ktransformers source code, the system recognizes five major model families with built-in configurations for their specific MoE implementations.
DeepSeek Model Support
KTransformers provides first-class support for the DeepSeek family of dense and MoE models, including production-ready configurations for recent variants.
DeepSeek-V3 and DeepSeek-R1 Variants
The model_registry.py file contains explicit entries for:
- DeepSeek-V3-0324
- DeepSeek-V3-2
- DeepSeek-R1-0528
- DeepSeek-V4-Flash
These entries define default parameters and kernel optimizations for each variant. The loader at kt-kernel/python/utils/loader.py (lines 1186-1189) includes specialized handling for DeepSeek-V4-Flash MXFP4 expert weights, enabling efficient quantization-aware inference.
Qwen Model Support (Qwen-2, Qwen-3, Qwen-3.5)
The Qwen MoE family is supported through string-based architecture detection in kt-kernel/python/sft/arch.py, which inspects model configurations to determine weight layout strategies.
Architecture Detection for Qwen MoE
The framework automatically detects Qwen variants using conditional checks:
if "Qwen2Moe" in arch or "Qwen3Moe" in arch or "Qwen3_5Moe" in arch:
# Apply Qwen-specific MoE routing
Supported configurations include:
- Qwen-2-Moe: Defined in
archive/ktransformers/models/configuration_qwen2_moe.py - Qwen-3-Moe: Defined in
archive/ktransformers/models/configuration_qwen3_moe.py - Qwen-3.5-Moe: Referenced in tests at
kt-kernel/test/per_commit/test_sft_shared_expert.py(lines 150-156) usingconfiguration.Qwen3_5MoeTextConfig
GLM-4 Model Support
KTransformers supports the GLM-4 family, specifically the MoE variants used in commercial deployments like GLM-4-5-Air.
The loader implementation in kt-kernel/python/utils/loader.py (lines 300-306) recognizes "GLM-4" model signatures and routes them to the appropriate handler. Model-specific configurations reside in archive/ktransformers/models/configuration_glm4_moe.py, which defines the expert routing dimensions and layer mappings required for the framework's optimized kernels.
MiniMax Model Support
The MiniMax family receives dedicated registry entries in kt-kernel/python/cli/utils/model_registry.py (lines 125-171).
Supported variants include:
- MiniMax-M2
- MiniMax-M2.1
These entries enable the framework to apply MiniMax-specific tensor parallelism strategies and expert sharding optimizations during inference.
Kimi-K2 Model Support
Kimi-K2-Thinking is officially supported through the model registry at kt-kernel/python/cli/utils/model_registry.py (lines 105-121).
The repository includes a dedicated conversion script for this model family:
python kt-kernel/scripts/convert_kimi_k2_fp8_to_bf16_cpu.py \
--input-fp8-hf-path /path/to/Kimi-K2-rawint4 \
--output-bf16-path /path/to/Kimi-K2-bf16
This utility converts FP8/Int4 checkpoints to CPU-friendly BF16 format, enabling local inference with KTransformers' AMX-optimized kernels.
Querying the Model Registry
You can programmatically inspect all supported models using the ModelRegistry class:
from kt_kernel.cli.utils.model_registry import ModelRegistry
# Initialise the built‑in registry
registry = ModelRegistry()
# List all supported model names
for name in registry._models:
print(name)
# Retrieve a specific model’s default parameters (e.g., DeepSeek‑V3‑0324)
deepseek_cfg = registry._models["DeepSeek-V3-0324"]
print(deepseek_cfg.default_params)
Architecture Detection and Loading
For dynamic model loading, KTransformers uses the architecture helper to map HuggingFace-style model names to internal layouts:
# Load a Qwen‑3.5 MoE model using the architecture helper
from kt_kernel.python.sft.arch import get_arch
arch = get_arch("Qwen3_5MoeForCausalLM")
print(arch) # → detects Qwen‑3.5 MoE layout
The get_arch function referenced above integrates with kt-kernel/python/utils/loader.py to instantiate the correct weight loader based on the detected architecture family, whether DeepSeek-style, Qwen-style, GLM-style, or MiniMax-style.
Summary
- KTransformers officially supports DeepSeek (V3-0324, V3-2, R1-0528, V4-Flash), Qwen (2/3/3.5 MoE), GLM-4, MiniMax (M2/M2.1), and Kimi-K2 model families.
- Model configurations are centralized in
kt-kernel/python/cli/utils/model_registry.pywith architecture detection handled bykt-kernel/python/sft/arch.py. - Each supported family has dedicated configuration files in
archive/ktransformers/models/defining MoE-specific parameters. - The loader system at
kt-kernel/python/utils/loader.pyautomatically applies family-specific optimizations including MXFP4/FP8 quantization for DeepSeek and BF16 conversion for Kimi-K2.
Frequently Asked Questions
How do I add a custom model to KTransformers?
You can extend support by adding entries to the ModelRegistry class in kt-kernel/python/cli/utils/model_registry.py and implementing the corresponding architecture detection logic in kt-kernel/python/sft/arch.py. Ensure your model follows MoE conventions compatible with the framework's expert parallelism kernels.
Does KTransformers support non-MoE models?
While the framework specializes in MoE architectures, the registry includes dense variants for some families. However, the optimized kernels (TRITON, FLASHINFER, AMX INT4/INT8) are specifically tuned for MoE expert routing patterns. Standard dense models may work but won't leverage the full optimization stack.
What quantization formats are supported for these models?
According to the source code, KTransformers supports MXFP4/FP8 (for DeepSeek-V4-Flash), INT4/INT8 via AMX instructions, and BF16 conversion for CPU inference. The specific format depends on the model family and target hardware backend.
Where can I find the configuration files for these models?
Architecture-specific configurations are located in archive/ktransformers/models/, including configuration_qwen2_moe.py, configuration_qwen3_moe.py, configuration_glm4_moe.py, and configuration_mini_max_moe.py. The active registry logic resides in kt-kernel/python/cli/utils/model_registry.py.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →