# How KTransformers Enables ROCm Support for AMD GPUs: A Complete Build Guide

> Learn how KTransformers brings ROCm support to AMD GPUs. This guide details the build process, enabling GPU acceleration for your AI models.

- Repository: [kvcache.ai/ktransformers](https://github.com/kvcache-ai/ktransformers)
- Tags: how-to-guide
- Published: 2026-07-26

---

**KTransformers supports AMD GPUs through a beta-level ROCm backend that uses compile-time flags, environment variables, and ROCm-specific PyTorch wheels to enable GPU acceleration while substituting unsupported Marlin kernels with generic `KLinearTorch` implementations.**

The `kvcache-ai/ktransformers` project provides official beta support for AMD GPU acceleration via the ROCm software stack. Unlike the mature CUDA backend, ROCm integration requires specific build-time configuration and runtime environment variables to activate the correct execution paths. This guide explains exactly how the library implements ROCm support based on the actual source code and build system configuration.

## Compile-Time Backend Selection

KTransformers enforces strict mutual exclusivity between GPU backends at the CMake level. In [`kt-kernel/CMakeLists.txt`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/CMakeLists.txt), the build system explicitly prevents simultaneous compilation of CUDA, SYCL, ROCm, MUSA, and MACA backends using a fatal error check:

```cmake
message(FATAL_ERROR "CUDA, SYCL, ROCm, MUSA, and MACA backends are mutually exclusive")

```

To select the ROCm backend, you must pass the CMake definition `-DKTRANSFORMERS_USE_ROCM=ON` during the build process. This flag triggers ROCm-specific preprocessor definitions and compiler paths throughout the `kt-kernel` library compilation.

## Environment Variables and Runtime Configuration

The library checks the environment variable `CPUINFER_USE_ROCM` at two critical stages: during C++ extension compilation and at runtime initialization.

In [`kt-kernel/setup.py`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/setup.py), the build script reads this variable using `_env_get_bool("CPUINFER_USE_ROCM", False)` to determine whether to inject ROCm compiler flags and link against ROCm libraries. When set to `1` or `true`, the build system prepends ROCm-specific include paths and disables CUDA-specific code paths.

At runtime, the same variable activates the ROCm execution branch within the kernel library, diverting tensor operations to AMD GPU-compatible implementations rather than the default CUDA kernels.

## Installing ROCm-Compatible PyTorch

KTransformers requires the official ROCm-enabled PyTorch distribution rather than the standard CUDA wheel. The recommended approach preserves the ROCm-compiled PyTorch binary throughout the build process:

```bash

# Install ROCm-compatible PyTorch (example for ROCm 6.2.4)

pip install torch torchvision torchaudio \
  --index-url https://download.pytorch.org/whl/rocm6.2.4

```

Installing the correct wheel beforehand is critical because the subsequent KTransformers build must use `--no-deps` to prevent pip from overwriting the ROCm PyTorch with a CUDA version.

## Building kt-kernel with ROCm Support

The `kt-kernel` package requires specific environment variables to compile correctly for AMD hardware. You must specify both the backend activation and target GPU architecture:

```bash

# Set ROCm build configuration

export CPUINFER_USE_ROCM=1                # Enable ROCm backend in setup.py

export PYTORCH_ROCM_ARCH=gfx936           # Target architecture (e.g., gfx936 for RX 7900 XTX)

export ROCM_PATH=/opt/rocm                # ROCm installation path (if non-standard)

# Build without isolation to preserve ROCm PyTorch

cd kt-kernel
pip install . --no-build-isolation --no-deps

```

The `--no-build-isolation` flag ensures the build process can access the previously installed ROCm PyTorch, while `--no-deps` prevents pip from resolving dependencies that might replace the ROCm wheel with a CUDA variant.

## ROCm-Specific Inference Configuration

KTransformers ships dedicated optimization YAML files for ROCm hardware located in `optimize/optimize_rules/rocm/`. These configurations replace the high-performance **Marlin** quantization kernels—which are not supported on ROCm—with the generic **KLinearTorch** implementation.

For example, [`optimize/optimize_rules/rocm/DeepSeek-V3-Chat.yaml`](https://github.com/kvcache-ai/ktransformers/blob/main/optimize/optimize_rules/rocm/DeepSeek-V3-Chat.yaml) maps linear layers to use `KLinearTorch` instead of `KLinearMarlin`. When running inference, explicitly point to these ROCm-tailored configurations:

```bash
python ktransformers/local_chat.py \
  --model_path deepseek-ai/DeepSeek-R1 \
  --gguf_path /path/to/gguf \
  --optimize_config_path ktransformers/optimize/optimize_rules/rocm/DeepSeek-V3-Chat.yaml \
  --cpu_infer $(($(nproc)+1))

```

Users with higher-VRAM GPUs can manually edit the YAML files to substitute `KLinearMarlin` with `KLinearTorch` for specific layer types if customizing beyond the provided defaults.

## Performance Considerations and Limitations

The current ROCm implementation lacks optimized Marlin Q8 kernels, resulting in modest performance compared to the CUDA backend. The `KLinearTorch` fallback provides compatibility but does not deliver the same inference speed as the quantized CUDA kernels.

Future releases aim to add native ROCm kernel implementations to close this performance gap. For production deployments requiring maximum throughput, CUDA hardware remains the recommended platform until ROCm kernel optimization is complete.

## Summary

- **Backend Exclusivity**: The [`kt-kernel/CMakeLists.txt`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/CMakeLists.txt) enforces single-backend compilation, requiring `-DKTRANSFORMERS_USE_ROCM=ON` to enable AMD support.
- **Environment Control**: Set `CPUINFER_USE_ROCM=1` before building and running to activate ROCm code paths in [`setup.py`](https://github.com/kvcache-ai/ktransformers/blob/main/setup.py) and the runtime library.
- **PyTorch Compatibility**: Install official ROCm PyTorch wheels before building to avoid CUDA/ROCm wheel conflicts during installation.
- **Architecture Targeting**: Use `PYTORCH_ROCM_ARCH` to specify your GPU's micro-architecture (e.g., `gfx936`) during compilation.
- **Kernel Substitution**: ROCm configurations in `optimize/optimize_rules/rocm/` replace unsupported Marlin kernels with `KLinearTorch` for functional compatibility.

## Frequently Asked Questions

### What AMD GPU architectures are supported by KTransformers?

KTransformers supports AMD GPUs that are compatible with ROCm-enabled PyTorch, including CDNA2 (MI200 series) and RDNA3 (RX 7000 series) architectures. You must specify your specific architecture code via the `PYTORCH_ROCM_ARCH` environment variable during compilation, such as `gfx90a` for MI200 accelerators or `gfx936` for consumer RDNA3 cards.

### Why does KTransformers use KLinearTorch instead of Marlin on ROCm?

The Marlin quantization kernels contain CUDA-specific optimizations that are not portable to AMD's ROCm stack. Since KTransformers cannot compile these kernels for AMD GPUs, the ROCm configuration files substitute `KLinearMarlin` with `KLinearTorch`, which uses standard PyTorch operations that ROCm can execute. This ensures functional inference at the cost of reduced performance compared to the optimized CUDA implementation.

### Can I mix CUDA and ROCm backends in the same KTransformers installation?

No. The [`kt-kernel/CMakeLists.txt`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/CMakeLists.txt) explicitly forbids mixing backends with a fatal error during configuration. You must choose exactly one GPU backend—CUDA, ROCm, SYCL, MUSA, or MACA—when compiling the kernel library. To switch between NVIDIA and AMD hardware, you need to rebuild the `kt-kernel` package from scratch with the appropriate backend flags.

### How do I prevent pip from overwriting my ROCm PyTorch during installation?

Use the `--no-build-isolation` and `--no-deps` flags when installing `kt-kernel`. The `--no-build-isolation` flag allows the build process to access your current environment (including the ROCm PyTorch), while `--no-deps` prevents pip from resolving dependencies that might pull in a CUDA-specific torch wheel. This preserves your ROCm installation while allowing KTransformers to compile against it.