# KTransformers Example Scripts: A Complete Guide to Kernel Testing and Fine-Tuning

> Explore KTransformers example scripts for kernel testing, Python API integration, and fine-tuning workflows. Discover efficient ways to test and optimize your models.

- Repository: [kvcache.ai/ktransformers](https://github.com/kvcache-ai/ktransformers)
- Tags: how-to-guide
- Published: 2026-07-21

---

**KTransformers ships with runnable example scripts in `kt-kernel/examples/` that demonstrate kernel-level testing, Python API integration, and end-to-end fine-tuning workflows.**

The KTransformers repository organizes its **ktransformers example scripts** into four distinct categories: low-level kernel tests, Python wrapper demonstrations, LLaMA-Factory fine-tuning pipelines, and model-specific serving tutorials. These scripts serve as canonical references for implementing optimized attention, mixture-of-experts (MoE), and rotary position embedding (RoPE) operations on hardware with AVX2, AMX, or CUDA support.

## Kernel-Level Test Scripts

The `kt-kernel/examples/` directory contains standalone tests that invoke C++/CUDA kernels directly from Python to verify correctness and benchmark performance.

### Attention Kernel Testing

In [`kt-kernel/examples/test_attention.py`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/examples/test_attention.py), the script demonstrates how to call the optimized attention kernel with dummy tensors:

```python
import torch
from ktransformers import attention

# Dummy Q, K, V tensors (batch=1, heads=8, seq_len=128, head_dim=64)

Q = torch.randn(1, 8, 128, 64, device="cuda")
K = torch.randn_like(Q)
V = torch.randn_like(Q)

# Compute attention scores and output

out = attention(q=Q, k=K, v=V, causal=True)

print("Attention output shape:", out.shape)   # → (1, 8, 128, 64)

```

This pattern validates the kernel implementation against PyTorch reference implementations and measures throughput on target hardware.

### MoE and Softmax Benchmarks

The repository includes dedicated scripts for testing specialized kernels. [`kt-kernel/examples/test_moe.py`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/examples/test_moe.py) exercises the mixture-of-experts implementation:

```python
import torch
from ktransformers import moe

# Simulate a MoE layer with 4 experts, each 256-dim

x = torch.randn(1, 1024, 256, device="cuda")
weights = torch.randn(4, 256, 256, device="cuda")
gate = torch.randn(1, 1024, 4, device="cuda")   # softmax-ed gating scores

# Forward pass through the MoE

y = moe(x, weights, gate)

print("MoE output shape:", y.shape)   # → (1, 1024, 256)

```

Similarly, [`kt-kernel/examples/test_softmax.py`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/examples/test_softmax.py) provides isolated benchmarks for the optimized softmax kernel, while [`kt-kernel/examples/test_rope.py`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/examples/test_rope.py) demonstrates rotary position embedding helpers:

```python
import torch
from ktransformers import rope

x = torch.randn(1, 8, 128, 64, device="cuda")
freqs = rope.precompute_freqs(dim=64, base=10000.0, seq_len=128)

# Apply RoPE in-place

rope.apply_rope_(x, freqs)

print("After RoPE:", x.shape)

```

## Python API Integration Examples

Beyond raw kernel tests, [`kt-kernel/examples/torch_attention.py`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/examples/torch_attention.py) illustrates how to integrate KTransformers kernels with PyTorch modules, including mixed-precision (FP16/BF16) configurations and CPU back-ends utilizing AVX2 or AMX instructions. These scripts assume the package is installed via `pip install ktransformers` and that hardware-specific libraries are available on the host.

## Fine-Tuning and Serving Examples

### LoRA Fine-Tuning with LLaMA-Factory

For supervised fine-tuning (SFT) workflows, KTransformers provides integration examples in `archive/kt-sft/csrc/ktransformers_ext/examples/`. The [`test_sft_moe.py`](https://github.com/kvcache-ai/ktransformers/blob/main/test_sft_moe.py) script demonstrates end-to-end LoRA training on MoE architectures using the KTransformers backend.

To launch fine-tuning, set the `USE_KT=1` environment variable when invoking LLaMA-Factory:

```bash
USE_KT=1 llamafactory-cli train \
    examples/train_lora/deepseek_v3_lora_sft_kt.yaml

```

The YAML configuration file ([`deepseek_v3_lora_sft_kt.yaml`](https://github.com/kvcache-ai/ktransformers/blob/main/deepseek_v3_lora_sft_kt.yaml)) specifies a **DeepSeek-V3** model with KTransformers acceleration enabled, configuring LoRA adapters for efficient parameter updates.

### Model Serving Workflows

The [`kt-kernel/README.md`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/README.md) documentation includes a complete serving example for the **Qwen3-30B-A3B** model:

```bash
USE_KT=1 python -m ktransformers.serve \
    --model Qwen3-30B-A3B \
    --backend native \
    --device cuda

```

This command launches a RESTful chat completion server that routes attention and feed-forward computations through KTransformers kernels, supporting the Native, AMX, and LLAMAFILE back-ends.

## Summary

- **Kernel tests** reside in `kt-kernel/examples/` and cover `attention()`, `moe()`, and `rope` operations with CUDA and CPU backends.
- **Python integration** examples demonstrate mixed-precision usage and hardware-specific optimizations (AVX2/AMX).
- **Fine-tuning scripts** in `archive/kt-sft/` enable LoRA training via LLaMA-Factory integration using the `USE_KT=1` flag.
- **Serving tutorials** provide copy-paste commands for deploying models like Qwen3-30B-A3B with optimized inference kernels.

## Frequently Asked Questions

### Where are the KTransformers example scripts located?

The primary examples live in `kt-kernel/examples/` for kernel testing and `archive/kt-sft/csrc/ktransformers_ext/examples/` for fine-tuning workflows. Documentation examples are embedded in [`kt-kernel/README.md`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/README.md).

### Do the examples require specific hardware?

Yes. The CUDA examples require an NVIDIA GPU, while CPU-optimized examples in [`torch_attention.py`](https://github.com/kvcache-ai/ktransformers/blob/main/torch_attention.py) assume AVX2 or AMX instruction set support. All examples require `pip install ktransformers` before execution.

### Can I use these scripts for production fine-tuning?

The [`test_sft_moe.py`](https://github.com/kvcache-ai/ktransformers/blob/main/test_sft_moe.py) and [`deepseek_v3_lora_sft_kt.yaml`](https://github.com/kvcache-ai/ktransformers/blob/main/deepseek_v3_lora_sft_kt.yaml) examples provide production-ready templates for LoRA fine-tuning. They are designed as starting points that you can adapt for custom datasets and model architectures within the LLaMA-Factory ecosystem.

### How do I benchmark kernel performance?

Use [`kt-kernel/examples/test_attention.py`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/examples/test_attention.py), [`test_softmax.py`](https://github.com/kvcache-ai/ktransformers/blob/main/test_softmax.py), or [`test_moe.py`](https://github.com/kvcache-ai/ktransformers/blob/main/test_moe.py) as benchmarks. These scripts measure execution time for core operations and can be modified to test different tensor shapes or batch sizes relevant to your deployment scenario.