# KTransformers Model Architectures: Supported Backends and Custom Integration Guide

> Explore KTransformers model architectures like Qwen2MoE, DeepSeek V2/V3, LLaMA, and GLM-4 MoE. Learn to add custom models by implementing wrappers and GGUF weight loading.

- Repository: [kvcache.ai/ktransformers](https://github.com/kvcache-ai/ktransformers)
- Tags: deep-dive
- Published: 2026-07-20

---

**KTransformers supports eight transformer architectures—including Qwen2MoE, DeepSeek V2/V3, LLaMA, and GLM-4 MoE—through modular wrappers that inherit from `BaseInjectedModule`, and you can add new models by implementing a config class, decoder layers, and a model wrapper that handles GGUF weight loading.**

KTransformers is an open-source inference engine that accelerates large language models through optimized GGUF loading and heterogeneous computing. Understanding which **KTransformers model architectures** are supported—and how to extend the framework—is essential for developers working with custom or bleeding-edge transformers.

## Currently Supported Model Architectures

KTransformers ships with ready-to-use transformer backends implemented as subclasses of `BaseInjectedModule`. Each backend wraps a Hugging Face-style model and handles GGUF weight loading, device placement, and optional per-layer pre-fill.

The following architectures are supported as of the latest release:

- **Qwen2MoE**: Implemented in `KQwen2MoeModel` (config: `Qwen2MoeConfig`) at `archive/ktransformers/operators/models.py:L185`
- **DeepSeek V2**: Implemented in `KDeepseekV2Model` (config: `DeepseekV2Config` via `DeepseekV2DecoderLayer`) at `archive/ktransformers/operators/models.py:L547`
- **LLaMA**: Implemented in `KLlamaModel` (config: `LlamaConfig`) at `archive/ktransformers/operators/models.py:L993`
- **Qwen3MoE**: Implemented in `KQwen3MoeModel` (config: `Qwen3MoeConfig`) at `archive/ktransformers/operators/models.py:L1466`
- **Qwen3Next**: Implemented in `KQwen3NextModel` (config: `Qwen3NextConfig`) in [`archive/ktransformers/operators/models.py`](https://github.com/kvcache-ai/ktransformers/blob/main/archive/ktransformers/operators/models.py)
- **SmallThinker**: Implemented in `KSmallThinkerModel` (config: `SmallThinkerConfig`) in [`archive/ktransformers/models/custom_modeling_smallthinker.py`](https://github.com/kvcache-ai/ktransformers/blob/main/archive/ktransformers/models/custom_modeling_smallthinker.py)
- **GLM-4 MoE**: Implemented in `KGlm4MoeModel` (config: `Glm4MoeConfig`) in [`archive/ktransformers/models/custom_modeling_glm4_moe.py`](https://github.com/kvcache-ai/ktransformers/blob/main/archive/ktransformers/models/custom_modeling_glm4_moe.py)
- **DeepSeek V3**: Implemented in `KDeepseekV3Model` (config: `DeepseekV3Config`) in [`archive/ktransformers/models/custom_modeling_deepseek_v3.py`](https://github.com/kvcache-ai/ktransformers/blob/main/archive/ktransformers/models/custom_modeling_deepseek_v3.py)

All implementations rely on the `GGUFLoader` class from [`ktransformers/util/custom_loader.py`](https://github.com/kvcache-ai/ktransformers/blob/main/ktransformers/util/custom_loader.py) for weight loading and `MOEArchConfig` from [`kt-kernel/python/sft/arch.py`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/python/sft/arch.py) for MoE-specific device placement.

## Core Architecture Design Pattern

Every supported model follows a consistent injection pattern centered on the `BaseInjectedModule` base class. This design allows KTransformers to intercept forward passes while maintaining compatibility with Hugging Face model signatures.

Each architecture implementation must:

1. **Inherit from `BaseInjectedModule`** to gain access to GGUF loading hooks and device management utilities.
2. **Accept a standard constructor signature** including `key` (model identifier), `gguf_loader` (instance of `GGUFLoader`), `config` (model-specific `PretrainedConfig`), `orig_module` (original Hugging Face module), and `device` placement string.
3. **Implement a `forward` method** that mirrors the original Hugging Face signature, accepting parameters such as `input_ids`, `attention_mask`, `position_ids`, `past_key_values`, `use_cache`, and `cache_position`.
4. **Optionally overload `load_layer_to`** to enable per-layer CPU/GPU movement for large-context scenarios or layer-wise pre-fill optimizations.

## How to Add Support for New Model Architectures

If your transformer architecture is not yet supported, you can integrate it with KTransformers by following these six implementation steps:

### 1. Create a Config Wrapper

Define a configuration class that inherits from `PretrainedConfig` and register it so `AutoConfig.from_pretrained` can locate it.

```python

# my_model/configuration_my_model.py

from transformers.configuration_utils import PretrainedConfig

class MyModelConfig(PretrainedConfig):
    model_type = "my_model"
    # Define architecture-specific hyperparameters

    hidden_size = 4096
    num_hidden_layers = 32

```

Export this class in your package's [`__init__.py`](https://github.com/kvcache-ai/ktransformers/blob/main/__init__.py) to ensure visibility.

### 2. Implement Decoder Layer(s)

Create modular decoder layers under `ktransformers/models` using standard PyTorch `nn.Module` components. Reuse existing building blocks such as `Attention`, `MLP`, and `RMSNorm` where possible. Ensure each layer supports `.to(device)` movement for heterogeneous execution.

### 3. Write the Model Wrapper

Subclass `BaseInjectedModule` to create your model backend (e.g., `KMyModel`). Mirror the original Hugging Face model's `forward` signature exactly to maintain compatibility with existing pipelines.

```python

# archive/ktransformers/operators/models.py

class KMyModel(BaseInjectedModule):
    def __init__(self, key: str, gguf_loader: GGUFLoader,
                 config: MyModelConfig, orig_module: nn.Module,
                 device: str = "cuda", **kwargs):
        super().__init__(key, gguf_loader, config, orig_module, device, **kwargs)
        # Model-specific initialization (e.g., per-layer thresholds)

    def forward(self, input_ids=None, attention_mask=None,
                position_ids=None, past_key_values=None,
                inputs_embeds=None, use_cache=None,
                output_attentions=None, output_hidden_states=None,
                return_dict=None, cache_position=None):
        # 1. Convert inputs_embeds if provided

        # 2. Build causal mask using _update_causal_mask

        # 3. Iterate over self.layers (decoder blocks)

        # 4. Apply final layer norm and return ModelOutputWithPast

        pass

```

Handle GGUF weight loading inside `forward` or initialization by accessing `self.embed_tokens` and `self.layers` through the provided `gguf_loader`.

### 4. Register the Model

Import your new class in the module's [`__init__.py`](https://github.com/kvcache-ai/ktransformers/blob/main/__init__.py) (e.g., [`ktransformers/__init__.py`](https://github.com/kvcache-ai/ktransformers/blob/main/ktransformers/__init__.py)). If using a factory function like `load_kt_model`, extend the model type mapping to associate your config's `model_type` with your wrapper class.

```python

# ktransformers/__init__.py

from .operators.models import KQwen2MoeModel, KDeepseekV2Model, KLlamaModel, KMyModel

```

### 5. Add Tests

Write unit tests under `kt-kernel/test/` that instantiate your model with a minimal GGUF checkpoint. Verify tensor shapes, cache handling, and optional per-layer pre-fill functionality. Run the repository's CI suite to detect regressions.

### 6. Update Documentation

Add your architecture to the supported models list and provide usage examples in the project documentation.

## Complete Integration Example

The following pattern demonstrates how to instantiate a custom model wrapper following the KTransformers convention:

```python
from ktransformers import KMyModel
from ktransformers.util.custom_loader import GGUFLoader
from transformers import AutoConfig

# Load GGUF weights

gguf_path = "/path/to/my_model.gguf"
loader = GGUFLoader(gguf_path)

# Load HuggingFace config and original module

config = AutoConfig.from_pretrained("my-org/my-model", trust_remote_code=True)
orig = MyOriginalHFModel.from_pretrained("my-org/my-model")

# Initialize KTransformers wrapper

model = KMyModel(
    key="my_model",
    gguf_loader=loader,
    config=config,
    orig_module=orig,
    device="cuda",
)

```

This follows the exact pattern used by built-in models such as `KLlamaModel` and `KDeepseekV2Model` in [`archive/ktransformers/operators/models.py`](https://github.com/kvcache-ai/ktransformers/blob/main/archive/ktransformers/operators/models.py).

## Summary

- **KTransformers supports eight architectures**: Qwen2MoE, DeepSeek V2/V3, LLaMA, Qwen3MoE, Qwen3Next, SmallThinker, and GLM-4 MoE.
- **All implementations inherit from `BaseInjectedModule`** and use `GGUFLoader` for weight management.
- **Adding new models requires six steps**: config wrapper, decoder layers, model wrapper, registration, testing, and documentation.
- **The `forward` method must match Hugging Face signatures** exactly to ensure pipeline compatibility.
- **File locations**: Core wrappers reside in [`archive/ktransformers/operators/models.py`](https://github.com/kvcache-ai/ktransformers/blob/main/archive/ktransformers/operators/models.py), while custom model definitions are in `archive/ktransformers/models/`.

## Frequently Asked Questions

### Which model architectures does KTransformers support out of the box?

KTransformers supports Qwen2MoE, DeepSeek V2, DeepSeek V3, LLaMA, Qwen3MoE, Qwen3Next, SmallThinker, and GLM-4 MoE. Each implementation resides in [`archive/ktransformers/operators/models.py`](https://github.com/kvcache-ai/ktransformers/blob/main/archive/ktransformers/operators/models.py) or dedicated custom modeling files such as [`custom_modeling_deepseek_v3.py`](https://github.com/kvcache-ai/ktransformers/blob/main/custom_modeling_deepseek_v3.py).

### What is the BaseInjectedModule class in KTransformers?

`BaseInjectedModule` is the abstract base class that all KTransformers model wrappers must inherit from. It provides standardized hooks for GGUF weight loading via `GGUFLoader`, device placement management, and optional per-layer pre-fill functionality through the `load_layer_to` method.

### How do I register a new model architecture in KTransformers?

Register your model by importing your wrapper class (e.g., `KMyModel`) in [`ktransformers/__init__.py`](https://github.com/kvcache-ai/ktransformers/blob/main/ktransformers/__init__.py) and extending any factory mapping that selects backends based on `model_type`. Ensure your config class is visible to `AutoConfig.from_pretrained` by exporting it in your package initialization.

### Can I use existing Hugging Face models with KTransformers?

Yes. KTransformers wraps existing Hugging Face models by accepting the original module as the `orig_module` parameter in the constructor. The wrapper delegates to the original model's weights while overriding the `forward` method to enable GGUF loading and optimized device placement.