How to Configure Multi-GPU Inference with Heterogeneous Expert Placement in KTransformers

KTransformers enables CPU-GPU heterogeneous inference for MoE models by generating masks that place frequently-used "hot" experts on GPU while keeping cold experts in CPU memory, allowing massive Mixture-of-Experts models to run across multiple GPUs with limited VRAM.

KTransformers provides a dedicated framework for deploying large Mixture-of-Experts (MoE) models on heterogeneous hardware. This guide demonstrates how to configure multi-GPU inference with heterogeneous expert placement in KTransformers, using mask-based routing to distribute computation across multiple accelerators while maintaining CPU-resident parameters via shared memory buffers.

The Seven-Step Deployment Workflow

Configuring distributed heterogeneous inference follows a predictable pipeline from statistics collection to runtime execution:

  1. Collect activation statistics – Gather a tensor of shape (num_layers, num_experts) representing expert usage frequency. This data drives the placement decision.
  2. Generate GPU expert masks – Use generate_gpu_experts_masks to select the top-K most active experts that will occupy GPU memory.
  3. Partition masks per GPU – Split the global mask into N subsets (where N is your GPU count), assigning specific expert indices to each device.
  4. Load and place model parameters – Load the model once per process, then call move_non_experts_to_gpu to migrate embeddings, norms, and routers to the target GPU while leaving selected experts on CPU.
  5. Initialize CPU buffers – Instantiate KExpertsCPUBuffer to create pinned-memory regions for CPU expert weights, enabling efficient cross-device transfer.
  6. Launch distributed processes – Start one inference process per GPU using torchrun or torch.distributed, passing the local mask and device ID.
  7. Configure the inference server – When using the SGLang integration, specify --kt-gpu-device and --kt-num-gpu-experts flags to notify the runtime of your heterogeneous setup.

Generating Expert Placement Masks

The generate_gpu_experts_masks utility in kt-kernel/python/experts_base.py (lines 21-72) converts activation frequencies into boolean masks. This function analyzes your (num_layers, num_experts) frequency tensor and returns a mask indicating which experts should reside on GPU.

from kt_kernel import generate_gpu_experts_masks

# act_freq shape: [num_layers, num_experts]

mask = generate_gpu_experts_masks(
    act_freq, 
    num_gpu_experts=8  # Total experts to place across all GPUs

)

The resulting mask is a torch.bool tensor on CPU that serves as the single source of truth for device placement across all distributed processes.

Moving Non-Expert Parameters to GPU

After generating masks, each process must migrate non-expert weights while preserving the expert parameters on CPU. The move_non_experts_to_gpu function in kt-kernel/python/sft/arch.py (lines 22-27) handles this automatically:

from kt_kernel.sft.arch import move_non_experts_to_gpu

# Move embeddings, norms, routers, and shared experts to cuda:0

# while keeping masked experts on CPU

move_non_experts_to_gpu(model, device="cuda:0")

This ensures that only the selected GPU experts and the current process's allocated subset consume device memory, while the MoE router and shared representations remain local for fast access.

Model-Agnostic MoE Architecture Detection

KTransformers automatically detects model-specific MoE configurations via get_moe_arch_config in kt-kernel/python/sft/arch.py (lines 62-99). This function builds a MOEArchConfig object that identifies which attributes contain the router gates, expert lists, and shared-expert modules for architectures including DeepSeek, Qwen, GLM, and Mixtral.

This abstraction allows the heterogeneous placement logic to work across different model families without manual configuration.

CPU Buffer Management for Multi-GPU Cooperation

The KExpertsCPUBuffer class in kt-kernel/python/experts_base.py (lines 42-75) provides pinned-memory buffers that are shared across processes. These buffers store CPU-resident expert weights and enable asynchronous data transfer to GPUs during the forward pass, eliminating allocation overhead during inference.

When an input token routes to a CPU-resident expert, the system streams weights from this buffer to the requesting GPU without blocking other devices.

Complete Multi-GPU Setup Example

The following implementation demonstrates a two-GPU setup using torchrun:

import os
import torch
from kt_kernel import generate_gpu_experts_masks
from kt_kernel.sft.arch import move_non_experts_to_gpu, get_moe_arch_config

# 1. Load activation frequencies (collected from calibration runs)

act_freq = torch.load("activation_freq.pt")  # shape: [L, E]

# 2. Generate global GPU mask

total_gpu_experts = 8
mask = generate_gpu_experts_masks(act_freq, num_gpu_experts=total_gpu_experts)

# 3. Split mask per GPU (simple even split)

rank = int(os.getenv("RANK", "0"))
gpu_mask = mask.clone()

if rank == 0:
    gpu_mask[:, total_gpu_experts // 2 :] = False
else:
    gpu_mask[:, : total_gpu_experts // 2] = False

# 4. Load model and move non-experts to target device

model = torch.load("my_moe_model.pt")
move_non_experts_to_gpu(model, device=f"cuda:{rank}")

# 5. Register mask for runtime consumption

model.register_buffer("gpu_expert_mask", gpu_mask)

# 6. Start inference (server or local loop)

# When using SGLang: pass --kt-gpu-device cuda:{rank} --kt-num-gpu-experts {total_gpu_experts}

Launching with SGLang Integration

For production deployments using the SGLang server, launch one process per GPU with the KTransformers-specific flags:


# On GPU 0

python -m sglang.server \
  --model-path /path/to/model \
  --kt-gpu-device cuda:0 \
  --kt-num-gpu-experts 8

# On GPU 1 (separate terminal or via torchrun)

python -m sglang.server \
  --model-path /path/to/model \
  --kt-gpu-device cuda:1 \
  --kt-num-gpu-experts 8

The --kt-num-gpu-experts flag tells the runtime the total number of experts residing on GPUs across all ranks, while --kt-gpu-device specifies the local CUDA device. The runtime automatically checks the gpu_expert_mask buffer to route tokens to the correct device.

Key Source Files

  • kt-kernel/python/experts_base.py – Contains generate_gpu_experts_masks for mask generation and KExpertsCPUBuffer for pinned memory management.
  • kt-kernel/python/sft/arch.py – Implements move_non_experts_to_gpu and get_moe_arch_config for architecture-agnostic parameter placement.
  • kt-kernel/test/test_generate_gpu_experts_masks.py – Reference test suite demonstrating mask generation patterns and edge case handling.
  • doc/en/kt-kernel/ – Tutorial documentation (e.g., DeepSeek-V3.2-sglang-tutorial.md) showing command-line integration examples.

Summary

Configuring multi-GPU inference with heterogeneous expert placement in KTransformers involves four core operations:

  • Generate placement masks using generate_gpu_experts_masks based on activation statistics to identify which experts benefit most from GPU residency.
  • Partition masks across distributed ranks so each GPU manages a distinct subset of hot experts.
  • Migrate parameters via move_non_experts_to_gpu to ensure embeddings, norms, and routers reside on the local device while MoE experts remain on CPU or specific GPUs.
  • Launch distributed processes with appropriate --kt-gpu-device and --kt-num-gpu-experts flags, leveraging KExpertsCPUBuffer for efficient CPU-to-GPU streaming of cold experts.

This architecture allows massive MoE models with hundreds of experts to run on consumer-grade multi-GPU setups by ensuring only the most frequently accessed parameters occupy limited VRAM.

Frequently Asked Questions

How does KTransformers decide which experts are "hot"?

KTransformers uses the generate_gpu_experts_masks function to select experts based on activation frequency tensors you provide. You must collect these statistics beforehand by running calibration data through your model, recording how often each expert in each layer is activated by the router. The function automatically selects the top-K most frequently used experts across all layers for GPU placement.

Can I use different partitioning strategies beyond simple splitting?

Yes. While the example demonstrates an even split between two GPUs, you can implement any custom heuristic in step three of the workflow. For instance, you could pack all hot experts from early layers on GPU 0 and later layers on GPU 1, or use a round-robin distribution. The mask is a standard boolean tensor that you can manipulate before registering it to the model via register_buffer("gpu_expert_mask", your_custom_mask).

What happens if a token routes to an expert not on the current GPU?

If a token routes to an expert marked False in the local GPU's mask, the inference engine automatically falls back to the CPU-resident expert via KExpertsCPUBuffer. The buffer provides pinned memory that enables asynchronous DMA transfer to the requesting GPU, minimizing latency. The router operates independently of placement, ensuring correct computation regardless of where parameters reside.

Which MoE architectures are supported for heterogeneous placement?

According to the get_moe_arch_config implementation in kt-kernel/python/sft/arch.py, KTransformers supports DeepSeek, Qwen, GLM, Mixtral, and other popular MoE architectures. The system automatically detects model-specific attributes (router names, expert lists, shared-expert configurations) by inspecting the HuggingFace config, requiring no manual architecture specification for supported models.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →