How KTransformers Handles MoE Expert Scheduling Between CPU and GPU
KTransformers uses a data-driven activation frequency scheduler to automatically place the most frequently used Mixture-of-Experts (MoE) experts on the GPU while keeping the rest on the CPU, enabling large models to run with limited VRAM.
The kvcache-ai/ktransformers repository implements a hybrid inference system that dynamically schedules MoE experts between system memory and GPU memory. This approach allows massive language models to exceed typical VRAM constraints by executing only the "hot" experts on the accelerator while leaving "cold" experts on the host. Understanding this scheduling mechanism is essential for optimizing throughput on resource-constrained hardware.
The Activation-Driven Scheduling Algorithm
The scheduler operates on a simple but effective premise: experts that are activated most frequently during inference should reside on the GPU for fast access, while rarely used experts can tolerate the higher latency of CPU execution.
Collecting Expert Activation Frequencies
When a model loads, KTransformers constructs an activation frequency table (activation_freq) with shape num_layers × num_experts. This tensor records how often each expert is selected during inference or training runs. The data in this table drives all subsequent placement decisions without requiring manual profiling or static heuristics.
Selecting Top-K Experts for GPU Placement
The core logic resides in generate_gpu_experts_masks located in kt-kernel/python/experts_base.py. The function flattens the activation frequency tensor, identifies the k highest values, and generates a Boolean mask indicating which experts should execute on the GPU.
# Implementation from kt-kernel/python/experts_base.py
flat_freq = activation_freq.view(-1).to(device="cpu")
_, top_indices = torch.topk(flat_freq, k=num_gpu_experts, largest=True, sorted=False)
gpu_experts_masks = torch.zeros(total_experts, dtype=torch.bool, device="cpu")
gpu_experts_masks[top_indices] = True
gpu_experts_masks = gpu_experts_masks.view(num_layers, num_experts_per_layer)
The mask creation handles edge cases by clamping the request: if num_gpu_experts exceeds the total expert count, it caps at the maximum, and negative values are treated as zero. This ensures the resulting mask always represents a valid subset of experts.
From Masks to Device Maps
Once the Boolean mask is generated, the system must translate these logical assignments into physical memory locations.
Building the Device Map
During model initialization in kt-kernel/python/sft/wrapper.py, the framework converts the Boolean mask into a device map dictionary. Each expert gets assigned to either torch.device("cuda") or torch.device("cpu") based on the mask value. This map is passed directly to KTMoEWrapper or AMXSFTMoEWrapper, which handle the actual parameter placement and forward pass orchestration.
The wrapper ensures that only experts marked True in the mask incur GPU memory allocation and compute costs. This selective placement is what allows KTransformers to load models with hundreds of experts on GPUs with limited VRAM.
Configuration and Runtime Execution
The scheduling system exposes user-configurable controls and manages the heterogeneous execution at inference time.
Environment-Based Configuration
Users control the number of GPU experts via the ACCELERATE_KT_NUM_GPU_EXPERTS environment variable, defined in kt-kernel/python/sft/config.py. The default value is 0, which places all experts on the CPU—a safe fallback for systems without dedicated accelerators. Setting this variable before model loading overrides the default and triggers the scheduling algorithm described above.
Hybrid Inference at Runtime
At inference time, the system copies GPU-selected expert parameters into pinned buffers managed by KExpertsCPUBuffer. Hot experts execute via optimized GPU kernels, while cold experts remain in system memory and run through CPU-optimized compute paths. This hybrid strategy minimizes data movement overhead by keeping frequently accessed weights resident in VRAM while allowing the total model size to exceed GPU capacity by an order of magnitude.
Complete Implementation Example
The following example demonstrates the full workflow from activation frequency analysis to device map construction:
import torch
from ktransformers import generate_gpu_experts_masks
# 1️⃣ Build a toy activation-frequency table (2 layers, 4 experts each)
activation_freq = torch.tensor([[0.1, 0.5, 0.3, 0.8],
[0.2, 0.4, 0.9, 0.1]])
# 2️⃣ Ask for 3 GPU experts (the 3 most active)
gpu_mask = generate_gpu_experts_masks(activation_freq, num_gpu_experts=3)
print(gpu_mask)
# → tensor([[False, True, False, True],
# [False, False, True, False]])
# 3️⃣ Use the mask to build a device map (simplified)
device_map = {}
for layer_idx, layer_mask in enumerate(gpu_mask):
for expert_idx, on_gpu in enumerate(layer_mask):
dev = "cuda" if on_gpu else "cpu"
device_map[(layer_idx, expert_idx)] = dev
# 4️⃣ Pass the map when loading a MoE model
# from ktransformers import KTMoEWrapper
# model = KTMoEWrapper.load_pretrained("my_moe_model", device_map=device_map)
This pattern is implemented across the key source files: kt-kernel/python/experts_base.py for mask generation, kt-kernel/python/sft/wrapper.py for device map construction, and kt-kernel/python/sft/config.py for configuration parsing. Unit tests in kt-kernel/test/test_generate_gpu_experts_masks.py validate the mask generation logic against various edge cases.
Summary
- KTransformers ranks MoE experts by activation frequency stored in a
num_layers × num_expertstensor and selects the top-K for GPU placement viagenerate_gpu_experts_masks. - The Boolean mask produced by the scheduler is converted into a device map in
kt-kernel/python/sft/wrapper.py, directing each expert tocudaorcpu. - Users configure GPU expert capacity through the
ACCELERATE_KT_NUM_GPU_EXPERTSenvironment variable, defaulting to 0 (CPU-only). - At runtime, hot experts execute in pinned GPU buffers while cold experts use CPU kernels, enabling hybrid inference that exceeds standalone GPU memory limits.
Frequently Asked Questions
How does KTransformers decide which experts to place on the GPU?
KTransformers analyzes historical activation frequencies collected in the activation_freq table. The generate_gpu_experts_masks function in kt-kernel/python/experts_base.py selects the num_gpu_experts most frequently activated entries using torch.topk, creating a mask that assigns these high-traffic experts to GPU execution while leaving others on the CPU.
What happens if I request more GPU experts than the model contains?
The scheduler clamps the request automatically. If num_gpu_experts exceeds the total number of available experts, the value is capped at the maximum expert count. Negative values are treated as zero, ensuring the mask always represents a valid subset that prevents out-of-bounds memory access.
Can I change the number of GPU experts without reloading the model?
No, the device map is constructed statically during model initialization in kt-kernel/python/sft/wrapper.py. To adjust GPU expert allocation, you must set a new value for ACCELERATE_KT_NUM_GPU_EXPERTS and reload the model, as the framework pre-allocates pinned buffers and kernel contexts based on the initial mask.
Where is the scheduling logic located in the source code?
The core scheduling algorithm lives in kt-kernel/python/experts_base.py within the generate_gpu_experts_masks function. Configuration handling resides in kt-kernel/python/sft/config.py, while the device map construction and wrapper initialization occur in kt-kernel/python/sft/wrapper.py. Public API exposure is handled in kt-kernel/python/__init__.py.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →