# How to Configure Multi-GPU Inference with Heterogeneous Expert Placement in KTransformers

> Learn how to configure multi-GPU inference with heterogeneous expert placement in KTransformers. Run massive MoE models across GPUs with limited VRAM by intelligently placing experts.

- Repository: [kvcache.ai/ktransformers](https://github.com/kvcache-ai/ktransformers)
- Tags: how-to-guide
- Published: 2026-07-26

---

**KTransformers enables CPU-GPU heterogeneous inference for MoE models by generating masks that place frequently-used "hot" experts on GPU while keeping cold experts in CPU memory, allowing massive Mixture-of-Experts models to run across multiple GPUs with limited VRAM.**

KTransformers provides a dedicated framework for deploying large Mixture-of-Experts (MoE) models on heterogeneous hardware. This guide demonstrates how to configure multi-GPU inference with heterogeneous expert placement in KTransformers, using mask-based routing to distribute computation across multiple accelerators while maintaining CPU-resident parameters via shared memory buffers.

## The Seven-Step Deployment Workflow

Configuring distributed heterogeneous inference follows a predictable pipeline from statistics collection to runtime execution:

1. **Collect activation statistics** – Gather a tensor of shape `(num_layers, num_experts)` representing expert usage frequency. This data drives the placement decision.
2. **Generate GPU expert masks** – Use `generate_gpu_experts_masks` to select the top-K most active experts that will occupy GPU memory.
3. **Partition masks per GPU** – Split the global mask into N subsets (where N is your GPU count), assigning specific expert indices to each device.
4. **Load and place model parameters** – Load the model once per process, then call `move_non_experts_to_gpu` to migrate embeddings, norms, and routers to the target GPU while leaving selected experts on CPU.
5. **Initialize CPU buffers** – Instantiate `KExpertsCPUBuffer` to create pinned-memory regions for CPU expert weights, enabling efficient cross-device transfer.
6. **Launch distributed processes** – Start one inference process per GPU using `torchrun` or `torch.distributed`, passing the local mask and device ID.
7. **Configure the inference server** – When using the SGLang integration, specify `--kt-gpu-device` and `--kt-num-gpu-experts` flags to notify the runtime of your heterogeneous setup.

## Generating Expert Placement Masks

The `generate_gpu_experts_masks` utility in [`kt-kernel/python/experts_base.py`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/python/experts_base.py) (lines 21-72) converts activation frequencies into boolean masks. This function analyzes your `(num_layers, num_experts)` frequency tensor and returns a mask indicating which experts should reside on GPU.

```python
from kt_kernel import generate_gpu_experts_masks

# act_freq shape: [num_layers, num_experts]

mask = generate_gpu_experts_masks(
    act_freq, 
    num_gpu_experts=8  # Total experts to place across all GPUs

)

```

The resulting mask is a `torch.bool` tensor on CPU that serves as the single source of truth for device placement across all distributed processes.

## Moving Non-Expert Parameters to GPU

After generating masks, each process must migrate non-expert weights while preserving the expert parameters on CPU. The `move_non_experts_to_gpu` function in [`kt-kernel/python/sft/arch.py`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/python/sft/arch.py) (lines 22-27) handles this automatically:

```python
from kt_kernel.sft.arch import move_non_experts_to_gpu

# Move embeddings, norms, routers, and shared experts to cuda:0

# while keeping masked experts on CPU

move_non_experts_to_gpu(model, device="cuda:0")

```

This ensures that only the selected GPU experts and the current process's allocated subset consume device memory, while the MoE router and shared representations remain local for fast access.

## Model-Agnostic MoE Architecture Detection

KTransformers automatically detects model-specific MoE configurations via `get_moe_arch_config` in [`kt-kernel/python/sft/arch.py`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/python/sft/arch.py) (lines 62-99). This function builds a `MOEArchConfig` object that identifies which attributes contain the router gates, expert lists, and shared-expert modules for architectures including DeepSeek, Qwen, GLM, and Mixtral.

This abstraction allows the heterogeneous placement logic to work across different model families without manual configuration.

## CPU Buffer Management for Multi-GPU Cooperation

The `KExpertsCPUBuffer` class in [`kt-kernel/python/experts_base.py`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/python/experts_base.py) (lines 42-75) provides pinned-memory buffers that are shared across processes. These buffers store CPU-resident expert weights and enable asynchronous data transfer to GPUs during the forward pass, eliminating allocation overhead during inference.

When an input token routes to a CPU-resident expert, the system streams weights from this buffer to the requesting GPU without blocking other devices.

## Complete Multi-GPU Setup Example

The following implementation demonstrates a two-GPU setup using `torchrun`:

```python
import os
import torch
from kt_kernel import generate_gpu_experts_masks
from kt_kernel.sft.arch import move_non_experts_to_gpu, get_moe_arch_config

# 1. Load activation frequencies (collected from calibration runs)

act_freq = torch.load("activation_freq.pt")  # shape: [L, E]

# 2. Generate global GPU mask

total_gpu_experts = 8
mask = generate_gpu_experts_masks(act_freq, num_gpu_experts=total_gpu_experts)

# 3. Split mask per GPU (simple even split)

rank = int(os.getenv("RANK", "0"))
gpu_mask = mask.clone()

if rank == 0:
    gpu_mask[:, total_gpu_experts // 2 :] = False
else:
    gpu_mask[:, : total_gpu_experts // 2] = False

# 4. Load model and move non-experts to target device

model = torch.load("my_moe_model.pt")
move_non_experts_to_gpu(model, device=f"cuda:{rank}")

# 5. Register mask for runtime consumption

model.register_buffer("gpu_expert_mask", gpu_mask)

# 6. Start inference (server or local loop)

# When using SGLang: pass --kt-gpu-device cuda:{rank} --kt-num-gpu-experts {total_gpu_experts}

```

## Launching with SGLang Integration

For production deployments using the SGLang server, launch one process per GPU with the KTransformers-specific flags:

```bash

# On GPU 0

python -m sglang.server \
  --model-path /path/to/model \
  --kt-gpu-device cuda:0 \
  --kt-num-gpu-experts 8

# On GPU 1 (separate terminal or via torchrun)

python -m sglang.server \
  --model-path /path/to/model \
  --kt-gpu-device cuda:1 \
  --kt-num-gpu-experts 8

```

The `--kt-num-gpu-experts` flag tells the runtime the total number of experts residing on GPUs across all ranks, while `--kt-gpu-device` specifies the local CUDA device. The runtime automatically checks the `gpu_expert_mask` buffer to route tokens to the correct device.

## Key Source Files

- **[`kt-kernel/python/experts_base.py`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/python/experts_base.py)** – Contains `generate_gpu_experts_masks` for mask generation and `KExpertsCPUBuffer` for pinned memory management.
- **[`kt-kernel/python/sft/arch.py`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/python/sft/arch.py)** – Implements `move_non_experts_to_gpu` and `get_moe_arch_config` for architecture-agnostic parameter placement.
- **[`kt-kernel/test/test_generate_gpu_experts_masks.py`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/test/test_generate_gpu_experts_masks.py)** – Reference test suite demonstrating mask generation patterns and edge case handling.
- **`doc/en/kt-kernel/`** – Tutorial documentation (e.g., *DeepSeek-V3.2-sglang-tutorial.md*) showing command-line integration examples.

## Summary

Configuring multi-GPU inference with heterogeneous expert placement in KTransformers involves four core operations:

- **Generate placement masks** using `generate_gpu_experts_masks` based on activation statistics to identify which experts benefit most from GPU residency.
- **Partition masks** across distributed ranks so each GPU manages a distinct subset of hot experts.
- **Migrate parameters** via `move_non_experts_to_gpu` to ensure embeddings, norms, and routers reside on the local device while MoE experts remain on CPU or specific GPUs.
- **Launch distributed processes** with appropriate `--kt-gpu-device` and `--kt-num-gpu-experts` flags, leveraging `KExpertsCPUBuffer` for efficient CPU-to-GPU streaming of cold experts.

This architecture allows massive MoE models with hundreds of experts to run on consumer-grade multi-GPU setups by ensuring only the most frequently accessed parameters occupy limited VRAM.

## Frequently Asked Questions

### How does KTransformers decide which experts are "hot"?

KTransformers uses the `generate_gpu_experts_masks` function to select experts based on activation frequency tensors you provide. You must collect these statistics beforehand by running calibration data through your model, recording how often each expert in each layer is activated by the router. The function automatically selects the top-K most frequently used experts across all layers for GPU placement.

### Can I use different partitioning strategies beyond simple splitting?

Yes. While the example demonstrates an even split between two GPUs, you can implement any custom heuristic in step three of the workflow. For instance, you could pack all hot experts from early layers on GPU 0 and later layers on GPU 1, or use a round-robin distribution. The mask is a standard boolean tensor that you can manipulate before registering it to the model via `register_buffer("gpu_expert_mask", your_custom_mask)`.

### What happens if a token routes to an expert not on the current GPU?

If a token routes to an expert marked `False` in the local GPU's mask, the inference engine automatically falls back to the CPU-resident expert via `KExpertsCPUBuffer`. The buffer provides pinned memory that enables asynchronous DMA transfer to the requesting GPU, minimizing latency. The router operates independently of placement, ensuring correct computation regardless of where parameters reside.

### Which MoE architectures are supported for heterogeneous placement?

According to the `get_moe_arch_config` implementation in [`kt-kernel/python/sft/arch.py`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/python/sft/arch.py), KTransformers supports DeepSeek, Qwen, GLM, Mixtral, and other popular MoE architectures. The system automatically detects model-specific attributes (router names, expert lists, shared-expert configurations) by inspecting the HuggingFace config, requiring no manual architecture specification for supported models.