# Capability Differences Between CUDA, DirectX, Metal, and CPU Backends in LuisaCompute

> Discover LuisaCompute's CUDA, DirectX, Metal, and CPU backend differences. Explore platform targets, code generation, ray-tracing APIs, and interoperability for optimal performance.

- Repository: [LuisaGroup/luisacompute](https://github.com/luisagroup/luisacompute)
- Tags: deep-dive
- Published: 2026-03-06

---

**LuisaCompute's four execution backends differ in target platforms, code-generation pipelines (PTX, HLSL, MSL, or LLVM IR), hardware-accelerated ray-tracing APIs, and cross-backend interoperability, with CUDA and DirectX offering the most advanced GPU features while the CPU backend provides a pure-software fallback.**

LuisaCompute is a high-performance compute framework that unifies GPU and CPU execution through a single DSL. Understanding the specific capability differences between its CUDA, DirectX 12, Metal, and CPU backends is essential for selecting the optimal target for rendering, AI inference, or general-purpose parallel workloads.

## Platform Targets and Runtime Environments

Each backend targets a distinct hardware and operating-system stack.

- **CUDA Backend**: Requires NVIDIA GPUs with CUDA-capable drivers. It is the primary choice for high-throughput computing on NVIDIA hardware, offering warp-level parallelism and direct access to NVIDIA RTX features.
- **DirectX 12 Backend**: Targets Windows 10/11 systems with DirectX 12.1-compatible GPUs. It is optimized for Windows-native graphics applications and provides access to DXR (DirectX Raytracing).
- **Metal Backend**: Supports macOS on both Apple Silicon and Intel GPUs. It leverages the Metal 3 API and Metal Ray-Tracing for hardware-accelerated rendering on Apple devices.
- **CPU Backend**: A pure-software fallback available on any x86-64 or ARM processor. It uses LLVM JIT/AOT compilation and is always available when GPU backends are not supported.

## Code Generation and Compilation Pipelines

The backends diverge significantly in how they transform LuisaCompute DSL kernels into executable machine code.

**CUDA** generates **PTX** or CUDA C++ source and compiles it at runtime using **NVRTC** or **NVCC**. The implementation resides in `src/backends/cuda/`, where the backend manages `CUmodule` and `CUfunction` objects.

**DirectX 12** generates **HLSL** source code and compiles it to DXIL using the **DXC** compiler. In `src/backends/dx/`, the backend creates `ID3D12PipelineState` objects and manages shader resources through DirectX 12's stateful pipeline.

**Metal** generates **Metal Shading Language (MSL)** and compiles it using the native Metal compiler. The implementation in `src/backends/metal/` handles `MTLLibrary` and `MTLFunction` creation, targeting the Apple GPU pipeline.

**CPU** generates **LLVM IR** and compiles it using the LLVM infrastructure for JIT or AOT execution. Located in `src/rust/`, this Rust-based backend produces standard functions executed via a `std::thread`-based thread pool.

## Runtime Architecture and Resource Management

Memory allocation and command scheduling architectures vary by backend.

- **CUDA**: Uses CUDA-specific allocators such as `cuMemAlloc` and memory pools. Commands are submitted to GPU-side streams with CUDA events for synchronization.
- **DirectX**: Employs the **D3D12MA** (DirectX 12 Memory Allocator) using a TLSF (Two-Level Segregated Fit) strategy. Work is recorded into DirectX command lists with explicit resource barriers.
- **Metal**: Utilizes **MTLHeap** for sub-allocation of GPU resources. Command encoding occurs through Metal command buffers and blit encoders.
- **CPU**: Relies on **mimalloc** or custom allocators. Execution follows a host-side task graph processed by an internal thread pool rather than GPU command queues.

## Ray Tracing and Hardware Acceleration

Hardware-accelerated ray tracing is supported on three backends, each using platform-specific APIs.

- **CUDA**: Implements ray tracing via **NVIDIA RTX** extensions, providing high-performance BVH traversal and intersection testing on RTX-capable hardware.
- **DirectX**: Uses **DXR** (DirectX 12 Raytracing) for hardware-accelerated ray tracing on compatible Windows GPUs.
- **Metal**: Leverages **Metal Ray-Tracing** (available in Metal 3) for accelerated ray tracing on Apple Silicon and recent AMD GPUs in Macs.
- **CPU**: Falls back to software-based ray tracing using **Embree**, Intel's high-performance ray tracing kernels. This provides correct results but at orders-of-magnitude lower performance than GPU alternatives.

## Interoperability and Cross-Backend Features

Only the CUDA and DirectX backends support resource sharing with external APIs.

- **CUDA Interop**: The CUDA backend supports direct interoperation with DirectX 12 resources via the `lc_dx_cuda_interop` path, as well as Vulkan-CUDA interop for cross-API workflows.
- **DirectX Interop**: The DirectX backend can share resources with CUDA through the same interop layer, enabling zero-copy data exchange between DXR ray tracing and CUDA compute kernels.
- **Metal and CPU**: Neither the Metal nor CPU backends currently expose native interoperability APIs for external graphics or compute frameworks.

## Build Configuration and Enable Flags

Each backend is conditionally compiled via CMake or XMake options.

- **CUDA**: Controlled by `LUISA_COMPUTE_ENABLE_CUDA` (default `ON`), implemented in `src/backends/cuda/`.
- **DirectX**: Controlled by `LUISA_COMPUTE_ENABLE_DX` (default `ON`), implemented in `src/backends/dx/`.
- **Metal**: Controlled by `LUISA_COMPUTE_ENABLE_METAL` (default `ON`), implemented in `src/backends/metal/`.
- **CPU**: Controlled by `LUISA_COMPUTE_ENABLE_CPU` (default `ON`), implemented in `src/rust/`.

## Usage Examples

The same DSL kernel works unchanged across all backends; only the device creation string differs.

```cpp
// Common kernel definition for all backends
Kernel2D fill = [&](ImageFloat img) noexcept {
    auto coord = dispatch_id().xy();
    auto size  = dispatch_size().xy();
    Float2 uv = make_float2(coord) / make_float2(size);
    img->write(coord, make_float4(uv, 0.5f, 1.0f));
};

```

**CUDA Backend**

```cpp
Context ctx{argv[0]};
Device dev = ctx.create_device("cuda");  // Selects CUDA backend
Stream s = dev.create_stream();
Image<float> img = dev.create_image<float>(PixelStorage::BYTE4, 1024, 1024);
auto shader = dev.compile(fill);
s << shader(img.view(0)).dispatch(1024, 1024) << synchronize();

```

**DirectX 12 Backend**

```cpp
Context ctx{argv[0]};
Device dev = ctx.create_device("dx");    // Selects DirectX backend
Stream s = dev.create_stream();
Image<float> img = dev.create_image<float>(PixelStorage::BYTE4, 1024, 1024);
auto shader = dev.compile(fill);
s << shader(img.view(0)).dispatch(1024, 1024) << synchronize();

```

**Metal Backend**

```cpp
Context ctx{argv[0]};
Device dev = ctx.create_device("metal"); // Selects Metal backend
Stream s = dev.create_stream();
Image<float> img = dev.create_image<float>(PixelStorage::BYTE4, 1024, 1024);
auto shader = dev.compile(fill);
s << shader(img.view(0)).dispatch(1024, 1024) << synchronize();

```

**CPU Backend**

```cpp
Context ctx{argv[0]};
Device dev = ctx.create_device("cpu");   // Selects CPU backend
Stream s = dev.create_stream();
Image<float> img = dev.create_image<float>(PixelStorage::BYTE4, 1024, 1024);
auto shader = dev.compile(fill);
s << shader(img.view(0)).dispatch(1024, 1024) << synchronize();

```

## Summary

- **CUDA** and **DirectX** provide the highest performance on their respective platforms, supporting hardware-accelerated ray tracing (NVIDIA RTX and DXR) and cross-API interoperability.
- **Metal** offers comparable capabilities on macOS, leveraging Metal Ray-Tracing for Apple Silicon GPUs.
- **CPU** serves as a universal fallback using LLVM-based compilation and Embree for ray tracing, ideal for debugging or headless environments without supported GPUs.
- Code generation differs fundamentally: PTX (CUDA), HLSL/DXIL (DirectX), MSL (Metal), and LLVM IR (CPU).
- Resource management uses platform-native allocators: `cuMemAlloc` (CUDA), D3D12MA (DirectX), MTLHeap (Metal), and mimalloc (CPU).

## Frequently Asked Questions

### Which LuisaCompute backend offers the best ray-tracing performance?

The **CUDA** backend utilizing NVIDIA RTX extensions typically delivers the highest ray-tracing performance on supported hardware, followed closely by the **DirectX** backend using DXR on Windows GPUs. The **Metal** backend provides optimized ray tracing for Apple Silicon via Metal Ray-Tracing, while the **CPU** backend relies on Embree and is significantly slower.

### Can I use LuisaCompute on a Mac without a dedicated GPU?

Yes. The **Metal** backend supports both Apple Silicon integrated GPUs and Intel-based Mac GPUs. Additionally, the **CPU** backend is always available as a fallback on any macOS system, compiling kernels to LLVM IR for execution on the host processor.

### How do I enable or disable specific backends when building LuisaCompute?

Use the CMake or XMake boolean flags defined in [`xmake.lua`](https://github.com/luisagroup/luisacompute/blob/main/xmake.lua) and [`CMakeLists.txt`](https://github.com/luisagroup/luisacompute/blob/main/CMakeLists.txt): `LUISA_COMPUTE_ENABLE_CUDA`, `LUISA_COMPUTE_ENABLE_DX`, `LUISA_COMPUTE_ENABLE_METAL`, and `LUISA_COMPUTE_ENABLE_CPU`. All default to `ON`, but you can set any to `OFF` to exclude that backend from the build.

### Does the CPU backend support the same features as the GPU backends?

The **CPU** backend supports the full LuisaCompute DSL and produces identical numerical results, but it lacks hardware-accelerated ray-tracing performance and GPU-specific interop capabilities. It uses software-based Embree for ray tracing instead of GPU RT cores, making it suitable for debugging and validation rather than production rendering.