How to Set Up NUMA-Aware Memory Management for KTransformers Performance

Configure KTransformers NUMA-aware memory management by running kt doctor to verify topology, using the --numa-nodes CLI flag or interactive wizard to create per-socket thread pools, and ensuring model weights follow the *.numa.{id}.weight naming convention.

KTransformers optimizes inference for CPU-heavy Mixture-of-Experts (MoE) models by exploiting Non-Uniform Memory Access (NUMA) architectures found in multi-socket servers. Proper NUMA-aware memory management ensures that expert weights reside in local memory to the CPU threads processing them, eliminating costly cross-socket traffic. This guide covers the complete configuration path from hardware detection in environment.py to runtime initialization in moe_kernel.py.

Detecting System NUMA Topology

KTransformers automatically discovers your hardware layout before allocating resources. In kt-kernel/python/utils/environment.py, the detect_cpu_info() function (line 289) reads /sys/devices/system/node/* to enumerate NUMA nodes and map physical CPU cores to each socket.

This detection drives all downstream configuration decisions. The function returns the node count and CPU affinity masks that the CLI and Python API use to validate user input. You can verify detection manually by checking your Linux sysfs:

ls /sys/devices/system/node

If you see node0, node1, and so on, your system exposes multiple NUMA domains that KTransformers can exploit.

Configuring NUMA-Aware Thread Pools

The runtime creates isolated thread pools bound to specific NUMA nodes to maintain memory locality. You can configure this through command-line flags or the interactive wizard.

CLI Configuration

The run and quant commands accept a --numa-nodes argument parsed in kt-kernel/python/cli/commands/run.py (line 49) and quant.py (line 68). Specify node IDs explicitly or let the system auto-detect:


# Automatic detection with 8 threads per NUMA node

kt run --model ./my_kt_model --cpu-threads-per-numa 8

# Manual mapping for a 4-socket system

kt run --model ./my_kt_model --numa-nodes 0 1 2 3 --cpu-threads-per-numa 6

Interactive Wizard Configuration

During the interactive setup wizard defined in kt-kernel/python/cli/utils/run_interactive.py (line 510), step 3 prompts for NUMA-specific settings:


# From run_interactive.py

console.print(Panel("[bold cyan]Step 3: NUMA and CPU Configuration[/bold cyan]"))
numa = Prompt.ask("NUMA Nodes (1 to {max})", default="1")
cpu_threads = Prompt.ask("CPU Threads per NUMA (1 to {max})", default="8")

The wizard validates your input against the topology detected earlier, preventing over-subscription of physical cores.

Loading NUMA-Sharded Expert Weights

For optimal performance, model weights must be pre-sharded across NUMA nodes following a specific naming convention. In kt-kernel/python/utils/loader.py (lines 84-86), the SafeTensorLoader.load_experts() method searches for tensor keys ending with .numa.{numa_id}.weight and .numa.{numa_id}.scale.

Each NUMA node loads only its local slice, ensuring that the massive parameter sets of MoE models never traverse the inter-socket link during inference:

from kt_kernel.python.utils.loader import SafeTensorLoader

loader = SafeTensorLoader("./weights")

# Load expert 3's weights for NUMA node 1 only

weight = loader.load_tensor("blk.0.ffn_up_exps.3.numa.1.weight", device="cpu")
scale  = loader.load_tensor("blk.0.ffn_up_exps.3.numa.1.scale", device="cpu")

Store your converted checkpoints with the .numa.{id}.weight suffix to enable this locality-aware loading.

Initializing NUMA-Aware MoE Wrappers

The Python wrappers forward NUMA configuration to the underlying C++ kernel (mtl::ThreadPool). In kt-kernel/python/experts_base.py (line 170), BaseMoEWrapper stores threadpool_count and an optional numa_nodes list. These values are passed through GeneralMoEWrapper in kt-kernel/python/utils/moe_kernel.py (line 52) into the MOEConfig struct.

When instantiating the wrapper programmatically, explicitly provide the NUMA mapping:

from kt_kernel.python.utils.moe_kernel import GeneralMoEWrapper

wrapper = GeneralMoEWrapper(
    layer_idx=0,
    num_experts=64,
    num_experts_per_tok=2,
    hidden_size=4096,
    moe_intermediate_size=1024,
    gpu_experts_mask=None,
    cpuinfer_threads=8,       # Threads per NUMA sub-pool

    threadpool_count=4,       # One pool per NUMA node

    weight_path="./weights",
    chunked_prefill_size=64,
    numa_nodes=[0, 1, 2, 3]   # Explicit NUMA node IDs

)

The C++ kernel uses this mapping to allocate per-NUMA buffers and pin thread pools to specific sockets via sched_setaffinity.

Verifying Your NUMA Configuration

Before running inference, confirm that KTransformers recognizes your hardware correctly using the diagnostic tool:

kt doctor

This command (supported by localized strings in kt-kernel/python/cli/i18n.py, line 127) outputs a NUMA Topology section:


NUMA Topology
  Nodes: 2
  Node0 CPUs: [0,1,2,3,4,5,6,7]
  Node1 CPUs: [8,9,10,11,12,13,14,15]

If the node count matches your physical sockets, the pipeline is ready for NUMA-aware execution.

Complete Configuration Examples

Combine the components above for end-to-end deployment on multi-socket servers.

Example: 2-Socket AMD EPYC via CLI

kt run --model /models/kt-deepseek \
       --numa-nodes 0 1 \
       --cpu-threads-per-numa 16 \
       --max-batch-size 1

Example: 4-Socket Intel Xeon via Python API

import torch
from kt_kernel.python.utils.moe_kernel import GeneralMoEWrapper

# 4 NUMA nodes, 64 total experts

wrapper = GeneralMoEWrapper(
    layer_idx=0,
    num_experts=64,
    num_experts_per_tok=2,
    hidden_size=7168,
    moe_intermediate_size=2048,
    cpuinfer_threads=12,
    threadpool_count=4,
    numa_nodes=[0, 1, 2, 3],
    weight_path="/data/numa_sharded_model"
)

# Map physical expert indices to logical slots

wrapper.load_weights(physical_to_logical_map_cpu=torch.arange(64))

Summary

  • Detection: detect_cpu_info() in environment.py automatically identifies NUMA nodes by reading /sys/devices/system/node.
  • Thread Pools: Use --numa-nodes in run.py or the interactive wizard to create threadpool_count pools aligned with sockets.
  • Weight Sharding: Store checkpoints with .numa.{id}.weight suffixes so SafeTensorLoader in loader.py distributes them locally.
  • Kernel Binding: GeneralMoEWrapper passes NUMA maps to the C++ kernel via MOEConfig in moe_kernel.py to enforce memory locality.
  • Verification: Run kt doctor to confirm topology detection before launching inference.

Frequently Asked Questions

How do I know if my server supports NUMA?

Run ls /sys/devices/system/node on Linux. If you see multiple node* directories (e.g., node0 and node1), your system has multiple NUMA domains. The kt doctor command in kvcache-ai/ktransformers will also display the detected node count and CPU mappings.

Can I use NUMA-aware mode with a single-socket CPU?

Yes, but it provides no benefit. Set --numa-nodes 0 or allow the interactive wizard to default to one node. The optimization becomes critical for dual-socket and quad-socket configurations where remote memory access latency exceeds 100 nanoseconds.

What happens if my weights are not sharded with the .numa.{id} suffix?

The SafeTensorLoader will fail to find NUMA-specific keys and may fall back to loading all weights on the first socket or raise a KeyError. Ensure your model conversion pipeline splits experts into separate files or keys following the expert_name.numa.0.weight pattern expected by loader.py.

Is there a performance penalty for cross-NUMA access in KTransformers?

Severe penalties exist without NUMA awareness. Accessing memory attached to a remote socket can reduce throughput by 30-50% on large MoE models due to limited inter-socket bandwidth. Binding threads and weights to the same NUMA node via --numa-nodes eliminates this bottleneck by keeping traffic on the local memory controller.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →