KTransformers Model Architectures: Supported Backends and Custom Integration Guide

KTransformers supports eight transformer architectures—including Qwen2MoE, DeepSeek V2/V3, LLaMA, and GLM-4 MoE—through modular wrappers that inherit from BaseInjectedModule, and you can add new models by implementing a config class, decoder layers, and a model wrapper that handles GGUF weight loading.

KTransformers is an open-source inference engine that accelerates large language models through optimized GGUF loading and heterogeneous computing. Understanding which KTransformers model architectures are supported—and how to extend the framework—is essential for developers working with custom or bleeding-edge transformers.

Currently Supported Model Architectures

KTransformers ships with ready-to-use transformer backends implemented as subclasses of BaseInjectedModule. Each backend wraps a Hugging Face-style model and handles GGUF weight loading, device placement, and optional per-layer pre-fill.

The following architectures are supported as of the latest release:

All implementations rely on the GGUFLoader class from ktransformers/util/custom_loader.py for weight loading and MOEArchConfig from kt-kernel/python/sft/arch.py for MoE-specific device placement.

Core Architecture Design Pattern

Every supported model follows a consistent injection pattern centered on the BaseInjectedModule base class. This design allows KTransformers to intercept forward passes while maintaining compatibility with Hugging Face model signatures.

Each architecture implementation must:

  1. Inherit from BaseInjectedModule to gain access to GGUF loading hooks and device management utilities.
  2. Accept a standard constructor signature including key (model identifier), gguf_loader (instance of GGUFLoader), config (model-specific PretrainedConfig), orig_module (original Hugging Face module), and device placement string.
  3. Implement a forward method that mirrors the original Hugging Face signature, accepting parameters such as input_ids, attention_mask, position_ids, past_key_values, use_cache, and cache_position.
  4. Optionally overload load_layer_to to enable per-layer CPU/GPU movement for large-context scenarios or layer-wise pre-fill optimizations.

How to Add Support for New Model Architectures

If your transformer architecture is not yet supported, you can integrate it with KTransformers by following these six implementation steps:

1. Create a Config Wrapper

Define a configuration class that inherits from PretrainedConfig and register it so AutoConfig.from_pretrained can locate it.


# my_model/configuration_my_model.py

from transformers.configuration_utils import PretrainedConfig

class MyModelConfig(PretrainedConfig):
    model_type = "my_model"
    # Define architecture-specific hyperparameters

    hidden_size = 4096
    num_hidden_layers = 32

Export this class in your package's __init__.py to ensure visibility.

2. Implement Decoder Layer(s)

Create modular decoder layers under ktransformers/models using standard PyTorch nn.Module components. Reuse existing building blocks such as Attention, MLP, and RMSNorm where possible. Ensure each layer supports .to(device) movement for heterogeneous execution.

3. Write the Model Wrapper

Subclass BaseInjectedModule to create your model backend (e.g., KMyModel). Mirror the original Hugging Face model's forward signature exactly to maintain compatibility with existing pipelines.


# archive/ktransformers/operators/models.py

class KMyModel(BaseInjectedModule):
    def __init__(self, key: str, gguf_loader: GGUFLoader,
                 config: MyModelConfig, orig_module: nn.Module,
                 device: str = "cuda", **kwargs):
        super().__init__(key, gguf_loader, config, orig_module, device, **kwargs)
        # Model-specific initialization (e.g., per-layer thresholds)

    def forward(self, input_ids=None, attention_mask=None,
                position_ids=None, past_key_values=None,
                inputs_embeds=None, use_cache=None,
                output_attentions=None, output_hidden_states=None,
                return_dict=None, cache_position=None):
        # 1. Convert inputs_embeds if provided

        # 2. Build causal mask using _update_causal_mask

        # 3. Iterate over self.layers (decoder blocks)

        # 4. Apply final layer norm and return ModelOutputWithPast

        pass

Handle GGUF weight loading inside forward or initialization by accessing self.embed_tokens and self.layers through the provided gguf_loader.

4. Register the Model

Import your new class in the module's __init__.py (e.g., ktransformers/__init__.py). If using a factory function like load_kt_model, extend the model type mapping to associate your config's model_type with your wrapper class.


# ktransformers/__init__.py

from .operators.models import KQwen2MoeModel, KDeepseekV2Model, KLlamaModel, KMyModel

5. Add Tests

Write unit tests under kt-kernel/test/ that instantiate your model with a minimal GGUF checkpoint. Verify tensor shapes, cache handling, and optional per-layer pre-fill functionality. Run the repository's CI suite to detect regressions.

6. Update Documentation

Add your architecture to the supported models list and provide usage examples in the project documentation.

Complete Integration Example

The following pattern demonstrates how to instantiate a custom model wrapper following the KTransformers convention:

from ktransformers import KMyModel
from ktransformers.util.custom_loader import GGUFLoader
from transformers import AutoConfig

# Load GGUF weights

gguf_path = "/path/to/my_model.gguf"
loader = GGUFLoader(gguf_path)

# Load HuggingFace config and original module

config = AutoConfig.from_pretrained("my-org/my-model", trust_remote_code=True)
orig = MyOriginalHFModel.from_pretrained("my-org/my-model")

# Initialize KTransformers wrapper

model = KMyModel(
    key="my_model",
    gguf_loader=loader,
    config=config,
    orig_module=orig,
    device="cuda",
)

This follows the exact pattern used by built-in models such as KLlamaModel and KDeepseekV2Model in archive/ktransformers/operators/models.py.

Summary

  • KTransformers supports eight architectures: Qwen2MoE, DeepSeek V2/V3, LLaMA, Qwen3MoE, Qwen3Next, SmallThinker, and GLM-4 MoE.
  • All implementations inherit from BaseInjectedModule and use GGUFLoader for weight management.
  • Adding new models requires six steps: config wrapper, decoder layers, model wrapper, registration, testing, and documentation.
  • The forward method must match Hugging Face signatures exactly to ensure pipeline compatibility.
  • File locations: Core wrappers reside in archive/ktransformers/operators/models.py, while custom model definitions are in archive/ktransformers/models/.

Frequently Asked Questions

Which model architectures does KTransformers support out of the box?

KTransformers supports Qwen2MoE, DeepSeek V2, DeepSeek V3, LLaMA, Qwen3MoE, Qwen3Next, SmallThinker, and GLM-4 MoE. Each implementation resides in archive/ktransformers/operators/models.py or dedicated custom modeling files such as custom_modeling_deepseek_v3.py.

What is the BaseInjectedModule class in KTransformers?

BaseInjectedModule is the abstract base class that all KTransformers model wrappers must inherit from. It provides standardized hooks for GGUF weight loading via GGUFLoader, device placement management, and optional per-layer pre-fill functionality through the load_layer_to method.

How do I register a new model architecture in KTransformers?

Register your model by importing your wrapper class (e.g., KMyModel) in ktransformers/__init__.py and extending any factory mapping that selects backends based on model_type. Ensure your config class is visible to AutoConfig.from_pretrained by exporting it in your package initialization.

Can I use existing Hugging Face models with KTransformers?

Yes. KTransformers wraps existing Hugging Face models by accepting the original module as the orig_module parameter in the constructor. The wrapper delegates to the original model's weights while overriding the forward method to enable GGUF loading and optimized device placement.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →