# How KTransformers Handles MoE Expert Scheduling Between CPU and GPU

> Discover how KTransformers optimizes MoE expert scheduling between CPU and GPU using an activation frequency scheduler. Run large models with limited VRAM by placing frequent experts on GPU.

- Repository: [kvcache.ai/ktransformers](https://github.com/kvcache-ai/ktransformers)
- Tags: internals
- Published: 2026-07-26

---

**KTransformers uses a data-driven activation frequency scheduler to automatically place the most frequently used Mixture-of-Experts (MoE) experts on the GPU while keeping the rest on the CPU, enabling large models to run with limited VRAM.**

The `kvcache-ai/ktransformers` repository implements a hybrid inference system that dynamically schedules MoE experts between system memory and GPU memory. This approach allows massive language models to exceed typical VRAM constraints by executing only the "hot" experts on the accelerator while leaving "cold" experts on the host. Understanding this scheduling mechanism is essential for optimizing throughput on resource-constrained hardware.

## The Activation-Driven Scheduling Algorithm

The scheduler operates on a simple but effective premise: experts that are activated most frequently during inference should reside on the GPU for fast access, while rarely used experts can tolerate the higher latency of CPU execution.

### Collecting Expert Activation Frequencies

When a model loads, KTransformers constructs an **activation frequency table** (`activation_freq`) with shape `num_layers × num_experts`. This tensor records how often each expert is selected during inference or training runs. The data in this table drives all subsequent placement decisions without requiring manual profiling or static heuristics.

### Selecting Top-K Experts for GPU Placement

The core logic resides in **`generate_gpu_experts_masks`** located in [`kt-kernel/python/experts_base.py`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/python/experts_base.py). The function flattens the activation frequency tensor, identifies the `k` highest values, and generates a Boolean mask indicating which experts should execute on the GPU.

```python

# Implementation from kt-kernel/python/experts_base.py

flat_freq = activation_freq.view(-1).to(device="cpu")
_, top_indices = torch.topk(flat_freq, k=num_gpu_experts, largest=True, sorted=False)
gpu_experts_masks = torch.zeros(total_experts, dtype=torch.bool, device="cpu")
gpu_experts_masks[top_indices] = True
gpu_experts_masks = gpu_experts_masks.view(num_layers, num_experts_per_layer)

```

The mask creation handles edge cases by clamping the request: if `num_gpu_experts` exceeds the total expert count, it caps at the maximum, and negative values are treated as zero. This ensures the resulting mask always represents a valid subset of experts.

## From Masks to Device Maps

Once the Boolean mask is generated, the system must translate these logical assignments into physical memory locations.

### Building the Device Map

During model initialization in [`kt-kernel/python/sft/wrapper.py`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/python/sft/wrapper.py), the framework converts the Boolean mask into a **device map** dictionary. Each expert gets assigned to either `torch.device("cuda")` or `torch.device("cpu")` based on the mask value. This map is passed directly to `KTMoEWrapper` or `AMXSFTMoEWrapper`, which handle the actual parameter placement and forward pass orchestration.

The wrapper ensures that only experts marked `True` in the mask incur GPU memory allocation and compute costs. This selective placement is what allows KTransformers to load models with hundreds of experts on GPUs with limited VRAM.

## Configuration and Runtime Execution

The scheduling system exposes user-configurable controls and manages the heterogeneous execution at inference time.

### Environment-Based Configuration

Users control the number of GPU experts via the **`ACCELERATE_KT_NUM_GPU_EXPERTS`** environment variable, defined in [`kt-kernel/python/sft/config.py`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/python/sft/config.py). The default value is `0`, which places all experts on the CPU—a safe fallback for systems without dedicated accelerators. Setting this variable before model loading overrides the default and triggers the scheduling algorithm described above.

### Hybrid Inference at Runtime

At inference time, the system copies GPU-selected expert parameters into pinned buffers managed by `KExpertsCPUBuffer`. Hot experts execute via optimized GPU kernels, while cold experts remain in system memory and run through CPU-optimized compute paths. This hybrid strategy minimizes data movement overhead by keeping frequently accessed weights resident in VRAM while allowing the total model size to exceed GPU capacity by an order of magnitude.

## Complete Implementation Example

The following example demonstrates the full workflow from activation frequency analysis to device map construction:

```python
import torch
from ktransformers import generate_gpu_experts_masks

# 1️⃣ Build a toy activation-frequency table (2 layers, 4 experts each)

activation_freq = torch.tensor([[0.1, 0.5, 0.3, 0.8],
                                [0.2, 0.4, 0.9, 0.1]])

# 2️⃣ Ask for 3 GPU experts (the 3 most active)

gpu_mask = generate_gpu_experts_masks(activation_freq, num_gpu_experts=3)
print(gpu_mask)

# → tensor([[False,  True, False,  True],

#           [False, False,  True, False]])

# 3️⃣ Use the mask to build a device map (simplified)

device_map = {}
for layer_idx, layer_mask in enumerate(gpu_mask):
    for expert_idx, on_gpu in enumerate(layer_mask):
        dev = "cuda" if on_gpu else "cpu"
        device_map[(layer_idx, expert_idx)] = dev

# 4️⃣ Pass the map when loading a MoE model

# from ktransformers import KTMoEWrapper

# model = KTMoEWrapper.load_pretrained("my_moe_model", device_map=device_map)

```

This pattern is implemented across the key source files: [`kt-kernel/python/experts_base.py`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/python/experts_base.py) for mask generation, [`kt-kernel/python/sft/wrapper.py`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/python/sft/wrapper.py) for device map construction, and [`kt-kernel/python/sft/config.py`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/python/sft/config.py) for configuration parsing. Unit tests in [`kt-kernel/test/test_generate_gpu_experts_masks.py`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/test/test_generate_gpu_experts_masks.py) validate the mask generation logic against various edge cases.

## Summary

- KTransformers ranks MoE experts by activation frequency stored in a `num_layers × num_experts` tensor and selects the top-K for GPU placement via **`generate_gpu_experts_masks`**.
- The Boolean mask produced by the scheduler is converted into a device map in **[`kt-kernel/python/sft/wrapper.py`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/python/sft/wrapper.py)**, directing each expert to `cuda` or `cpu`.
- Users configure GPU expert capacity through the **`ACCELERATE_KT_NUM_GPU_EXPERTS`** environment variable, defaulting to 0 (CPU-only).
- At runtime, hot experts execute in pinned GPU buffers while cold experts use CPU kernels, enabling hybrid inference that exceeds standalone GPU memory limits.

## Frequently Asked Questions

### How does KTransformers decide which experts to place on the GPU?

KTransformers analyzes historical activation frequencies collected in the `activation_freq` table. The **`generate_gpu_experts_masks`** function in [`kt-kernel/python/experts_base.py`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/python/experts_base.py) selects the `num_gpu_experts` most frequently activated entries using `torch.topk`, creating a mask that assigns these high-traffic experts to GPU execution while leaving others on the CPU.

### What happens if I request more GPU experts than the model contains?

The scheduler clamps the request automatically. If `num_gpu_experts` exceeds the total number of available experts, the value is capped at the maximum expert count. Negative values are treated as zero, ensuring the mask always represents a valid subset that prevents out-of-bounds memory access.

### Can I change the number of GPU experts without reloading the model?

No, the device map is constructed statically during model initialization in [`kt-kernel/python/sft/wrapper.py`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/python/sft/wrapper.py). To adjust GPU expert allocation, you must set a new value for **`ACCELERATE_KT_NUM_GPU_EXPERTS`** and reload the model, as the framework pre-allocates pinned buffers and kernel contexts based on the initial mask.

### Where is the scheduling logic located in the source code?

The core scheduling algorithm lives in **[`kt-kernel/python/experts_base.py`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/python/experts_base.py)** within the `generate_gpu_experts_masks` function. Configuration handling resides in **[`kt-kernel/python/sft/config.py`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/python/sft/config.py)**, while the device map construction and wrapper initialization occur in **[`kt-kernel/python/sft/wrapper.py`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/python/sft/wrapper.py)**. Public API exposure is handled in [`kt-kernel/python/__init__.py`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/python/__init__.py).