# How to Perform DPO Fine-Tuning with KTransformers: A Complete LLaMA-Factory Guide

> Learn DPO fine-tuning with KTransformers and LLaMA Factory. Optimize massive MoE models on consumer GPUs by leveraging CPU CPU/AMX accelerators. Get started today.

- Repository: [kvcache.ai/ktransformers](https://github.com/kvcache-ai/ktransformers)
- Tags: how-to-guide
- Published: 2026-07-26

---

**KTransformers enables Direct Preference Optimization (DPO) fine-tuning of massive MoE models like DeepSeek-V3-671B on 2-4 RTX 4090 GPUs by offloading heavy attention and expert computations to CPU/AMX accelerators while LLaMA-Factory manages the LoRA training loop.**

The `kvcache-ai/ktransformers` repository provides a high-performance backend that integrates with the LLaMA-Factory framework to make DPO fine-tuning accessible for extremely large models. By enabling the `use_kt` flag and supplying a KTransformers **optimize rule**, you can train 671B parameter models with minimal GPU memory (as low as 6GB for 14B models) through a hybrid execution pipeline.

## Architecture Overview

KTransformers acts as an accelerated kernel replacement layer within the LLaMA-Factory training ecosystem. When `use_kt: true` is set in your configuration, KTransformers intercepts the model's **Attention** and **MoE (Mixture of Experts)** layers, routing computation according to a user-defined optimization rule.

The architecture follows this division of labor:

- **LLaMA-Factory** – Orchestrates data loading, LoRA adapter injection, learning rate scheduling, logging, and DPO-specific loss calculations (`pref_loss: sigmoid`, `pref_beta`).
- **KTransformers** – Replaces default HuggingFace kernels with optimized implementations. Attention layers typically execute on GPU while expert computations run on CPU or Intel AMX, as defined in the `kt_optimize_rule` YAML.
- **Optimize Rule** – A YAML configuration that matches layer names via regex (e.g., `^model\\.layers\\..*\\.mlp\\.experts$`) and specifies backend implementations (`AMXInt8`, `AMXBF16`, or `llamafile`), device placement (`prefill_device`, `generate_device`), and batch processing (`chunk_size`).

## Why DPO Works with Hybrid Execution

Direct Preference Optimization is a **LoRA-based** fine-tuning method that keeps the base model frozen while updating only a small set of adapter weights. Because the heavy attention and MoE computations dominate runtime rather than gradient updates, KTransformers' CPU-driven expert-parallel pipeline dramatically reduces GPU memory pressure.

This design specifically benefits DPO training because:
- The base model remains static, allowing KTransformers to optimize the frozen computation graph without tracking gradients through the full parameter set.
- DPO hyperparameters (`stage: dpo`, `pref_beta: 0.1`) are handled entirely by LLaMA-Factory; KTransformers only provides accelerated operators and does not need to understand preference optimization logic.

## Environment Setup

Create a dedicated conda environment and install the required dependencies for the KTransformers SFT (Supervised Fine-Tuning) package alongside LLaMA-Factory.

```bash
conda create -n Kllama python=3.12
conda install -y -c conda-forge libstdcxx-ng gcc_impl_linux-64
conda install -y -c nvidia/label/cuda-12.8.0 cuda-runtime

# Install LLaMA-Factory

git clone --depth 1 https://github.com/hiyouga/LLaMA-Factory.git
cd LLaMA-Factory
pip install -e ".[torch,metrics]" --no-build-isolation

# Install KTransformers with SFT support

pip install "ktransformers[sft]"

# Optional: FlashAttention for optimized attention kernels

pip install flash-attn --no-build-isolation

# Optional: Custom flash-infer backend

git clone https://github.com/kvcache-ai/custom_flashinfer.git
pip install custom_flashinfer/

```

## Configuring DPO Training

DPO fine-tuning requires two configuration files: a training configuration for LLaMA-Factory and an optimization rule for KTransformers.

### Training Configuration YAML

Create a YAML file (e.g., [`examples/train_lora/deepseek2_lora_dpo_kt.yaml`](https://github.com/kvcache-ai/ktransformers/blob/main/examples/train_lora/deepseek2_lora_dpo_kt.yaml)) specifying the DPO stage and KTransformers integration:

```yaml
model:
  model_name_or_path: deepseek-ai/DeepSeek-V2-Lite
  trust_remote_code: true

method:
  stage: dpo
  do_train: true
  finetuning_type: lora
  lora_rank: 8
  lora_target: all
  pref_beta: 0.1
  pref_loss: sigmoid   # DPO loss type

ktransformers:
  use_kt: true
  kt_optimize_rule: examples/kt_optimize_rules/DeepSeek-V2-Lite-Chat-sft-amx.yaml
  cpu_infer: 64
  chunk_size: 8192

```

### Optimization Rules (`kt_optimize_rule`)

The optimization rule defines which operators are swapped for KTransformers kernels. Store this in [`examples/kt_optimize_rules/DeepSeek-V2-Lite-Chat-sft-amx.yaml`](https://github.com/kvcache-ai/ktransformers/blob/main/examples/kt_optimize_rules/DeepSeek-V2-Lite-Chat-sft-amx.yaml):

```yaml
- match:
    name: "^model\\.layers\\..*\\.mlp\\.experts$"
  replace:
    class: ktransformers.operators.experts.KTransformersExperts
    kwargs:
      prefill_device: "cpu"
      prefill_op: "KExpertsTorch"
      generate_device: "cpu"
      generate_op: "KSFTExpertsCPU"
      out_device: "cuda"
      backend: "AMXInt8"

```

Key parameters include:
- **`match.name`** – Regex pattern identifying layers to replace.
- **`backend`** – Compute backend (`AMXInt8`, `AMXBF16`, or `llamafile`).
- **`prefill_device`**/`**generate_device`** – Where to run prefill vs. generation phases (typically "cpu" for experts, "cuda" for attention).

## Running DPO Fine-Tuning

Execute training with the `USE_KT=1` environment variable to enable KTransformers acceleration:

```bash
USE_KT=1 llamafactory-cli train examples/train_lora/deepseek2_lora_dpo_kt.yaml

```

This command invokes the LLaMA-Factory CLI with KTransformers kernel substitution active. According to the source code in [`ktransformers/sft/metrics_utils/constants.py`](https://github.com/kvcache-ai/ktransformers/blob/main/ktransformers/sft/metrics_utils/constants.py), the framework recognizes `"DPO"` as a supported training stage, ensuring compatibility with standard DPO loss calculations while KTransformers handles the underlying tensor operations.

## Deploying the Fine-Tuned Model

After training completes, load your LoRA-adapted model for inference using the chat interface:

```bash
llamafactory-cli chat examples/inference/deepseek2_lora_dpo_kt.yaml

```

Ensure your inference configuration points to the same `kt_optimize_rule` used during training to maintain consistent kernel acceleration and memory placement.

## Summary

- **KTransformers enables DPO fine-tuning** of massive MoE models (e.g., DeepSeek-V3-671B) on modest hardware (2-4 RTX 4090) through CPU/GPU hybrid execution.
- **Integration is controlled by `use_kt: true`** and an `kt_optimize_rule` YAML that specifies which layers to offload to CPU/AMX.
- **DPO-specific parameters** (`pref_beta`, `pref_loss`) are managed by LLaMA-Factory, while KTransformers provides accelerated kernels for the frozen base model.
- **Key files** referenced include [`DPO_tutorial.md`](https://github.com/kvcache-ai/ktransformers/blob/main/DPO_tutorial.md) and [`KTransformers-Fine-Tuning_User-Guide.md`](https://github.com/kvcache-ai/ktransformers/blob/main/KTransformers-Fine-Tuning_User-Guide.md) in the `doc/en/SFT/` directory.
- **Installation** requires the `ktransformers[sft]` package and the `USE_KT=1` environment variable during training.

## Frequently Asked Questions

### What hardware is required for DPO fine-tuning with KTransformers?

You can fine-tune models like DeepSeek-V3-671B using just 2-4 RTX 4090 GPUs. KTransformers offloads the expert layers to CPU memory, reducing GPU VRAM requirements significantly—down to approximately 6GB for a 14B parameter model. An Intel CPU with AMX support provides additional acceleration when using the `AMXInt8` or `AMXBF16` backends.

### How does KTransformers integrate with the LLaMA-Factory training loop?

KTransformers operates as a kernel replacement layer. When you set `use_kt: true` in your YAML configuration, KTransformers intercepts calls to attention and MoE layers, replacing them with optimized implementations. The LLaMA-Factory framework remains responsible for the training orchestration, LoRA management, and DPO loss calculation, as defined in [`kt-sft/ktransformers/sft/metrics_utils/constants.py`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-sft/ktransformers/sft/metrics_utils/constants.py).

### What is the purpose of the `kt_optimize_rule` YAML file?

The optimization rule file maps specific model layers (identified by regex patterns) to KTransformers operators and specifies their execution device. It controls whether experts run on CPU (`KSFTExpertsCPU`), GPU, or AMX, and defines the quantization backend (`AMXInt8`, etc.). This rule file is essential for configuring the hybrid execution that makes large-model fine-tuning feasible on consumer hardware.

### Can I use KTransformers for fine-tuning methods other than DPO?

Yes. While this guide focuses on DPO fine-tuning, KTransformers supports various LoRA-based fine-tuning stages available in LLaMA-Factory. The [`constants.py`](https://github.com/kvcache-ai/ktransformers/blob/main/constants.py) file in the SFT package defines supported stages, and the architecture supports any method where the base model remains frozen while adapter weights are updated, including SFT and other preference optimization variants.