# How to Set Up NUMA-Aware Memory Management for KTransformers Performance

> Boost KTransformers performance with NUMA-aware memory management. Learn to configure topology, thread pools, and model weights for optimal speed.

- Repository: [kvcache.ai/ktransformers](https://github.com/kvcache-ai/ktransformers)
- Tags: performance
- Published: 2026-07-26

---

**Configure KTransformers NUMA-aware memory management by running `kt doctor` to verify topology, using the `--numa-nodes` CLI flag or interactive wizard to create per-socket thread pools, and ensuring model weights follow the `*.numa.{id}.weight` naming convention.**

KTransformers optimizes inference for CPU-heavy Mixture-of-Experts (MoE) models by exploiting Non-Uniform Memory Access (NUMA) architectures found in multi-socket servers. Proper **NUMA-aware memory management** ensures that expert weights reside in local memory to the CPU threads processing them, eliminating costly cross-socket traffic. This guide covers the complete configuration path from hardware detection in [`environment.py`](https://github.com/kvcache-ai/ktransformers/blob/main/environment.py) to runtime initialization in [`moe_kernel.py`](https://github.com/kvcache-ai/ktransformers/blob/main/moe_kernel.py).

## Detecting System NUMA Topology

KTransformers automatically discovers your hardware layout before allocating resources. In [`kt-kernel/python/utils/environment.py`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/python/utils/environment.py), the `detect_cpu_info()` function (line 289) reads `/sys/devices/system/node/*` to enumerate NUMA nodes and map physical CPU cores to each socket.

This detection drives all downstream configuration decisions. The function returns the node count and CPU affinity masks that the CLI and Python API use to validate user input. You can verify detection manually by checking your Linux sysfs:

```bash
ls /sys/devices/system/node

```

If you see `node0`, `node1`, and so on, your system exposes multiple NUMA domains that KTransformers can exploit.

## Configuring NUMA-Aware Thread Pools

The runtime creates isolated thread pools bound to specific NUMA nodes to maintain memory locality. You can configure this through command-line flags or the interactive wizard.

### CLI Configuration

The `run` and `quant` commands accept a `--numa-nodes` argument parsed in [`kt-kernel/python/cli/commands/run.py`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/python/cli/commands/run.py) (line 49) and [`quant.py`](https://github.com/kvcache-ai/ktransformers/blob/main/quant.py) (line 68). Specify node IDs explicitly or let the system auto-detect:

```bash

# Automatic detection with 8 threads per NUMA node

kt run --model ./my_kt_model --cpu-threads-per-numa 8

# Manual mapping for a 4-socket system

kt run --model ./my_kt_model --numa-nodes 0 1 2 3 --cpu-threads-per-numa 6

```

### Interactive Wizard Configuration

During the interactive setup wizard defined in [`kt-kernel/python/cli/utils/run_interactive.py`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/python/cli/utils/run_interactive.py) (line 510), step 3 prompts for NUMA-specific settings:

```python

# From run_interactive.py

console.print(Panel("[bold cyan]Step 3: NUMA and CPU Configuration[/bold cyan]"))
numa = Prompt.ask("NUMA Nodes (1 to {max})", default="1")
cpu_threads = Prompt.ask("CPU Threads per NUMA (1 to {max})", default="8")

```

The wizard validates your input against the topology detected earlier, preventing over-subscription of physical cores.

## Loading NUMA-Sharded Expert Weights

For optimal performance, model weights must be pre-sharded across NUMA nodes following a specific naming convention. In [`kt-kernel/python/utils/loader.py`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/python/utils/loader.py) (lines 84-86), the `SafeTensorLoader.load_experts()` method searches for tensor keys ending with `.numa.{numa_id}.weight` and `.numa.{numa_id}.scale`.

Each NUMA node loads only its local slice, ensuring that the massive parameter sets of MoE models never traverse the inter-socket link during inference:

```python
from kt_kernel.python.utils.loader import SafeTensorLoader

loader = SafeTensorLoader("./weights")

# Load expert 3's weights for NUMA node 1 only

weight = loader.load_tensor("blk.0.ffn_up_exps.3.numa.1.weight", device="cpu")
scale  = loader.load_tensor("blk.0.ffn_up_exps.3.numa.1.scale", device="cpu")

```

Store your converted checkpoints with the `.numa.{id}.weight` suffix to enable this locality-aware loading.

## Initializing NUMA-Aware MoE Wrappers

The Python wrappers forward NUMA configuration to the underlying C++ kernel (`mtl::ThreadPool`). In [`kt-kernel/python/experts_base.py`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/python/experts_base.py) (line 170), `BaseMoEWrapper` stores `threadpool_count` and an optional `numa_nodes` list. These values are passed through `GeneralMoEWrapper` in [`kt-kernel/python/utils/moe_kernel.py`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/python/utils/moe_kernel.py) (line 52) into the `MOEConfig` struct.

When instantiating the wrapper programmatically, explicitly provide the NUMA mapping:

```python
from kt_kernel.python.utils.moe_kernel import GeneralMoEWrapper

wrapper = GeneralMoEWrapper(
    layer_idx=0,
    num_experts=64,
    num_experts_per_tok=2,
    hidden_size=4096,
    moe_intermediate_size=1024,
    gpu_experts_mask=None,
    cpuinfer_threads=8,       # Threads per NUMA sub-pool

    threadpool_count=4,       # One pool per NUMA node

    weight_path="./weights",
    chunked_prefill_size=64,
    numa_nodes=[0, 1, 2, 3]   # Explicit NUMA node IDs

)

```

The C++ kernel uses this mapping to allocate per-NUMA buffers and pin thread pools to specific sockets via `sched_setaffinity`.

## Verifying Your NUMA Configuration

Before running inference, confirm that KTransformers recognizes your hardware correctly using the diagnostic tool:

```bash
kt doctor

```

This command (supported by localized strings in [`kt-kernel/python/cli/i18n.py`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/python/cli/i18n.py), line 127) outputs a **NUMA Topology** section:

```

NUMA Topology
  Nodes: 2
  Node0 CPUs: [0,1,2,3,4,5,6,7]
  Node1 CPUs: [8,9,10,11,12,13,14,15]

```

If the node count matches your physical sockets, the pipeline is ready for NUMA-aware execution.

## Complete Configuration Examples

Combine the components above for end-to-end deployment on multi-socket servers.

**Example: 2-Socket AMD EPYC via CLI**

```bash
kt run --model /models/kt-deepseek \
       --numa-nodes 0 1 \
       --cpu-threads-per-numa 16 \
       --max-batch-size 1

```

**Example: 4-Socket Intel Xeon via Python API**

```python
import torch
from kt_kernel.python.utils.moe_kernel import GeneralMoEWrapper

# 4 NUMA nodes, 64 total experts

wrapper = GeneralMoEWrapper(
    layer_idx=0,
    num_experts=64,
    num_experts_per_tok=2,
    hidden_size=7168,
    moe_intermediate_size=2048,
    cpuinfer_threads=12,
    threadpool_count=4,
    numa_nodes=[0, 1, 2, 3],
    weight_path="/data/numa_sharded_model"
)

# Map physical expert indices to logical slots

wrapper.load_weights(physical_to_logical_map_cpu=torch.arange(64))

```

## Summary

- **Detection**: `detect_cpu_info()` in [`environment.py`](https://github.com/kvcache-ai/ktransformers/blob/main/environment.py) automatically identifies NUMA nodes by reading `/sys/devices/system/node`.
- **Thread Pools**: Use `--numa-nodes` in [`run.py`](https://github.com/kvcache-ai/ktransformers/blob/main/run.py) or the interactive wizard to create `threadpool_count` pools aligned with sockets.
- **Weight Sharding**: Store checkpoints with `.numa.{id}.weight` suffixes so `SafeTensorLoader` in [`loader.py`](https://github.com/kvcache-ai/ktransformers/blob/main/loader.py) distributes them locally.
- **Kernel Binding**: `GeneralMoEWrapper` passes NUMA maps to the C++ kernel via `MOEConfig` in [`moe_kernel.py`](https://github.com/kvcache-ai/ktransformers/blob/main/moe_kernel.py) to enforce memory locality.
- **Verification**: Run `kt doctor` to confirm topology detection before launching inference.

## Frequently Asked Questions

### How do I know if my server supports NUMA?

Run `ls /sys/devices/system/node` on Linux. If you see multiple `node*` directories (e.g., `node0` and `node1`), your system has multiple NUMA domains. The `kt doctor` command in kvcache-ai/ktransformers will also display the detected node count and CPU mappings.

### Can I use NUMA-aware mode with a single-socket CPU?

Yes, but it provides no benefit. Set `--numa-nodes 0` or allow the interactive wizard to default to one node. The optimization becomes critical for dual-socket and quad-socket configurations where remote memory access latency exceeds 100 nanoseconds.

### What happens if my weights are not sharded with the `.numa.{id}` suffix?

The `SafeTensorLoader` will fail to find NUMA-specific keys and may fall back to loading all weights on the first socket or raise a `KeyError`. Ensure your model conversion pipeline splits experts into separate files or keys following the `expert_name.numa.0.weight` pattern expected by [`loader.py`](https://github.com/kvcache-ai/ktransformers/blob/main/loader.py).

### Is there a performance penalty for cross-NUMA access in KTransformers?

Severe penalties exist without NUMA awareness. Accessing memory attached to a remote socket can reduce throughput by 30-50% on large MoE models due to limited inter-socket bandwidth. Binding threads and weights to the same NUMA node via `--numa-nodes` eliminates this bottleneck by keeping traffic on the local memory controller.