KTransformers Example Scripts: A Complete Guide to Kernel Testing and Fine-Tuning

KTransformers ships with runnable example scripts in kt-kernel/examples/ that demonstrate kernel-level testing, Python API integration, and end-to-end fine-tuning workflows.

The KTransformers repository organizes its ktransformers example scripts into four distinct categories: low-level kernel tests, Python wrapper demonstrations, LLaMA-Factory fine-tuning pipelines, and model-specific serving tutorials. These scripts serve as canonical references for implementing optimized attention, mixture-of-experts (MoE), and rotary position embedding (RoPE) operations on hardware with AVX2, AMX, or CUDA support.

Kernel-Level Test Scripts

The kt-kernel/examples/ directory contains standalone tests that invoke C++/CUDA kernels directly from Python to verify correctness and benchmark performance.

Attention Kernel Testing

In kt-kernel/examples/test_attention.py, the script demonstrates how to call the optimized attention kernel with dummy tensors:

import torch
from ktransformers import attention

# Dummy Q, K, V tensors (batch=1, heads=8, seq_len=128, head_dim=64)

Q = torch.randn(1, 8, 128, 64, device="cuda")
K = torch.randn_like(Q)
V = torch.randn_like(Q)

# Compute attention scores and output

out = attention(q=Q, k=K, v=V, causal=True)

print("Attention output shape:", out.shape)   # → (1, 8, 128, 64)

This pattern validates the kernel implementation against PyTorch reference implementations and measures throughput on target hardware.

MoE and Softmax Benchmarks

The repository includes dedicated scripts for testing specialized kernels. kt-kernel/examples/test_moe.py exercises the mixture-of-experts implementation:

import torch
from ktransformers import moe

# Simulate a MoE layer with 4 experts, each 256-dim

x = torch.randn(1, 1024, 256, device="cuda")
weights = torch.randn(4, 256, 256, device="cuda")
gate = torch.randn(1, 1024, 4, device="cuda")   # softmax-ed gating scores

# Forward pass through the MoE

y = moe(x, weights, gate)

print("MoE output shape:", y.shape)   # → (1, 1024, 256)

Similarly, kt-kernel/examples/test_softmax.py provides isolated benchmarks for the optimized softmax kernel, while kt-kernel/examples/test_rope.py demonstrates rotary position embedding helpers:

import torch
from ktransformers import rope

x = torch.randn(1, 8, 128, 64, device="cuda")
freqs = rope.precompute_freqs(dim=64, base=10000.0, seq_len=128)

# Apply RoPE in-place

rope.apply_rope_(x, freqs)

print("After RoPE:", x.shape)

Python API Integration Examples

Beyond raw kernel tests, kt-kernel/examples/torch_attention.py illustrates how to integrate KTransformers kernels with PyTorch modules, including mixed-precision (FP16/BF16) configurations and CPU back-ends utilizing AVX2 or AMX instructions. These scripts assume the package is installed via pip install ktransformers and that hardware-specific libraries are available on the host.

Fine-Tuning and Serving Examples

LoRA Fine-Tuning with LLaMA-Factory

For supervised fine-tuning (SFT) workflows, KTransformers provides integration examples in archive/kt-sft/csrc/ktransformers_ext/examples/. The test_sft_moe.py script demonstrates end-to-end LoRA training on MoE architectures using the KTransformers backend.

To launch fine-tuning, set the USE_KT=1 environment variable when invoking LLaMA-Factory:

USE_KT=1 llamafactory-cli train \
    examples/train_lora/deepseek_v3_lora_sft_kt.yaml

The YAML configuration file (deepseek_v3_lora_sft_kt.yaml) specifies a DeepSeek-V3 model with KTransformers acceleration enabled, configuring LoRA adapters for efficient parameter updates.

Model Serving Workflows

The kt-kernel/README.md documentation includes a complete serving example for the Qwen3-30B-A3B model:

USE_KT=1 python -m ktransformers.serve \
    --model Qwen3-30B-A3B \
    --backend native \
    --device cuda

This command launches a RESTful chat completion server that routes attention and feed-forward computations through KTransformers kernels, supporting the Native, AMX, and LLAMAFILE back-ends.

Summary

  • Kernel tests reside in kt-kernel/examples/ and cover attention(), moe(), and rope operations with CUDA and CPU backends.
  • Python integration examples demonstrate mixed-precision usage and hardware-specific optimizations (AVX2/AMX).
  • Fine-tuning scripts in archive/kt-sft/ enable LoRA training via LLaMA-Factory integration using the USE_KT=1 flag.
  • Serving tutorials provide copy-paste commands for deploying models like Qwen3-30B-A3B with optimized inference kernels.

Frequently Asked Questions

Where are the KTransformers example scripts located?

The primary examples live in kt-kernel/examples/ for kernel testing and archive/kt-sft/csrc/ktransformers_ext/examples/ for fine-tuning workflows. Documentation examples are embedded in kt-kernel/README.md.

Do the examples require specific hardware?

Yes. The CUDA examples require an NVIDIA GPU, while CPU-optimized examples in torch_attention.py assume AVX2 or AMX instruction set support. All examples require pip install ktransformers before execution.

Can I use these scripts for production fine-tuning?

The test_sft_moe.py and deepseek_v3_lora_sft_kt.yaml examples provide production-ready templates for LoRA fine-tuning. They are designed as starting points that you can adapt for custom datasets and model architectures within the LLaMA-Factory ecosystem.

How do I benchmark kernel performance?

Use kt-kernel/examples/test_attention.py, test_softmax.py, or test_moe.py as benchmarks. These scripts measure execution time for core operations and can be modified to test different tensor shapes or batch sizes relevant to your deployment scenario.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →