Capability Differences Between CUDA, DirectX, Metal, and CPU Backends in LuisaCompute

LuisaCompute's four execution backends differ in target platforms, code-generation pipelines (PTX, HLSL, MSL, or LLVM IR), hardware-accelerated ray-tracing APIs, and cross-backend interoperability, with CUDA and DirectX offering the most advanced GPU features while the CPU backend provides a pure-software fallback.

LuisaCompute is a high-performance compute framework that unifies GPU and CPU execution through a single DSL. Understanding the specific capability differences between its CUDA, DirectX 12, Metal, and CPU backends is essential for selecting the optimal target for rendering, AI inference, or general-purpose parallel workloads.

Platform Targets and Runtime Environments

Each backend targets a distinct hardware and operating-system stack.

  • CUDA Backend: Requires NVIDIA GPUs with CUDA-capable drivers. It is the primary choice for high-throughput computing on NVIDIA hardware, offering warp-level parallelism and direct access to NVIDIA RTX features.
  • DirectX 12 Backend: Targets Windows 10/11 systems with DirectX 12.1-compatible GPUs. It is optimized for Windows-native graphics applications and provides access to DXR (DirectX Raytracing).
  • Metal Backend: Supports macOS on both Apple Silicon and Intel GPUs. It leverages the Metal 3 API and Metal Ray-Tracing for hardware-accelerated rendering on Apple devices.
  • CPU Backend: A pure-software fallback available on any x86-64 or ARM processor. It uses LLVM JIT/AOT compilation and is always available when GPU backends are not supported.

Code Generation and Compilation Pipelines

The backends diverge significantly in how they transform LuisaCompute DSL kernels into executable machine code.

CUDA generates PTX or CUDA C++ source and compiles it at runtime using NVRTC or NVCC. The implementation resides in src/backends/cuda/, where the backend manages CUmodule and CUfunction objects.

DirectX 12 generates HLSL source code and compiles it to DXIL using the DXC compiler. In src/backends/dx/, the backend creates ID3D12PipelineState objects and manages shader resources through DirectX 12's stateful pipeline.

Metal generates Metal Shading Language (MSL) and compiles it using the native Metal compiler. The implementation in src/backends/metal/ handles MTLLibrary and MTLFunction creation, targeting the Apple GPU pipeline.

CPU generates LLVM IR and compiles it using the LLVM infrastructure for JIT or AOT execution. Located in src/rust/, this Rust-based backend produces standard functions executed via a std::thread-based thread pool.

Runtime Architecture and Resource Management

Memory allocation and command scheduling architectures vary by backend.

  • CUDA: Uses CUDA-specific allocators such as cuMemAlloc and memory pools. Commands are submitted to GPU-side streams with CUDA events for synchronization.
  • DirectX: Employs the D3D12MA (DirectX 12 Memory Allocator) using a TLSF (Two-Level Segregated Fit) strategy. Work is recorded into DirectX command lists with explicit resource barriers.
  • Metal: Utilizes MTLHeap for sub-allocation of GPU resources. Command encoding occurs through Metal command buffers and blit encoders.
  • CPU: Relies on mimalloc or custom allocators. Execution follows a host-side task graph processed by an internal thread pool rather than GPU command queues.

Ray Tracing and Hardware Acceleration

Hardware-accelerated ray tracing is supported on three backends, each using platform-specific APIs.

  • CUDA: Implements ray tracing via NVIDIA RTX extensions, providing high-performance BVH traversal and intersection testing on RTX-capable hardware.
  • DirectX: Uses DXR (DirectX 12 Raytracing) for hardware-accelerated ray tracing on compatible Windows GPUs.
  • Metal: Leverages Metal Ray-Tracing (available in Metal 3) for accelerated ray tracing on Apple Silicon and recent AMD GPUs in Macs.
  • CPU: Falls back to software-based ray tracing using Embree, Intel's high-performance ray tracing kernels. This provides correct results but at orders-of-magnitude lower performance than GPU alternatives.

Interoperability and Cross-Backend Features

Only the CUDA and DirectX backends support resource sharing with external APIs.

  • CUDA Interop: The CUDA backend supports direct interoperation with DirectX 12 resources via the lc_dx_cuda_interop path, as well as Vulkan-CUDA interop for cross-API workflows.
  • DirectX Interop: The DirectX backend can share resources with CUDA through the same interop layer, enabling zero-copy data exchange between DXR ray tracing and CUDA compute kernels.
  • Metal and CPU: Neither the Metal nor CPU backends currently expose native interoperability APIs for external graphics or compute frameworks.

Build Configuration and Enable Flags

Each backend is conditionally compiled via CMake or XMake options.

  • CUDA: Controlled by LUISA_COMPUTE_ENABLE_CUDA (default ON), implemented in src/backends/cuda/.
  • DirectX: Controlled by LUISA_COMPUTE_ENABLE_DX (default ON), implemented in src/backends/dx/.
  • Metal: Controlled by LUISA_COMPUTE_ENABLE_METAL (default ON), implemented in src/backends/metal/.
  • CPU: Controlled by LUISA_COMPUTE_ENABLE_CPU (default ON), implemented in src/rust/.

Usage Examples

The same DSL kernel works unchanged across all backends; only the device creation string differs.

// Common kernel definition for all backends
Kernel2D fill = [&](ImageFloat img) noexcept {
    auto coord = dispatch_id().xy();
    auto size  = dispatch_size().xy();
    Float2 uv = make_float2(coord) / make_float2(size);
    img->write(coord, make_float4(uv, 0.5f, 1.0f));
};

CUDA Backend

Context ctx{argv[0]};
Device dev = ctx.create_device("cuda");  // Selects CUDA backend
Stream s = dev.create_stream();
Image<float> img = dev.create_image<float>(PixelStorage::BYTE4, 1024, 1024);
auto shader = dev.compile(fill);
s << shader(img.view(0)).dispatch(1024, 1024) << synchronize();

DirectX 12 Backend

Context ctx{argv[0]};
Device dev = ctx.create_device("dx");    // Selects DirectX backend
Stream s = dev.create_stream();
Image<float> img = dev.create_image<float>(PixelStorage::BYTE4, 1024, 1024);
auto shader = dev.compile(fill);
s << shader(img.view(0)).dispatch(1024, 1024) << synchronize();

Metal Backend

Context ctx{argv[0]};
Device dev = ctx.create_device("metal"); // Selects Metal backend
Stream s = dev.create_stream();
Image<float> img = dev.create_image<float>(PixelStorage::BYTE4, 1024, 1024);
auto shader = dev.compile(fill);
s << shader(img.view(0)).dispatch(1024, 1024) << synchronize();

CPU Backend

Context ctx{argv[0]};
Device dev = ctx.create_device("cpu");   // Selects CPU backend
Stream s = dev.create_stream();
Image<float> img = dev.create_image<float>(PixelStorage::BYTE4, 1024, 1024);
auto shader = dev.compile(fill);
s << shader(img.view(0)).dispatch(1024, 1024) << synchronize();

Summary

  • CUDA and DirectX provide the highest performance on their respective platforms, supporting hardware-accelerated ray tracing (NVIDIA RTX and DXR) and cross-API interoperability.
  • Metal offers comparable capabilities on macOS, leveraging Metal Ray-Tracing for Apple Silicon GPUs.
  • CPU serves as a universal fallback using LLVM-based compilation and Embree for ray tracing, ideal for debugging or headless environments without supported GPUs.
  • Code generation differs fundamentally: PTX (CUDA), HLSL/DXIL (DirectX), MSL (Metal), and LLVM IR (CPU).
  • Resource management uses platform-native allocators: cuMemAlloc (CUDA), D3D12MA (DirectX), MTLHeap (Metal), and mimalloc (CPU).

Frequently Asked Questions

Which LuisaCompute backend offers the best ray-tracing performance?

The CUDA backend utilizing NVIDIA RTX extensions typically delivers the highest ray-tracing performance on supported hardware, followed closely by the DirectX backend using DXR on Windows GPUs. The Metal backend provides optimized ray tracing for Apple Silicon via Metal Ray-Tracing, while the CPU backend relies on Embree and is significantly slower.

Can I use LuisaCompute on a Mac without a dedicated GPU?

Yes. The Metal backend supports both Apple Silicon integrated GPUs and Intel-based Mac GPUs. Additionally, the CPU backend is always available as a fallback on any macOS system, compiling kernels to LLVM IR for execution on the host processor.

How do I enable or disable specific backends when building LuisaCompute?

Use the CMake or XMake boolean flags defined in xmake.lua and CMakeLists.txt: LUISA_COMPUTE_ENABLE_CUDA, LUISA_COMPUTE_ENABLE_DX, LUISA_COMPUTE_ENABLE_METAL, and LUISA_COMPUTE_ENABLE_CPU. All default to ON, but you can set any to OFF to exclude that backend from the build.

Does the CPU backend support the same features as the GPU backends?

The CPU backend supports the full LuisaCompute DSL and produces identical numerical results, but it lacks hardware-accelerated ray-tracing performance and GPU-specific interop capabilities. It uses software-based Embree for ray tracing instead of GPU RT cores, making it suitable for debugging and validation rather than production rendering.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →