# How the DeviceInterface Abstraction Enables Platform-Specific Backend Implementations in Luisa Compute

> Discover how the DeviceInterface abstraction in Luisa Compute isolates the runtime from GPU specifics. Enable platform-specific backends like CUDA, Metal, and DirectX with a unified API.

- Repository: [LuisaGroup/luisacompute](https://github.com/luisagroup/luisacompute)
- Tags: internals
- Published: 2026-03-06

---

**The `DeviceInterface` abstraction in Luisa Compute defines a pure-virtual contract that isolates the runtime from GPU-specific details, allowing backends like CUDA, Metal, and DirectX to implement platform-specific logic while presenting a unified API to user code.**

The `luisagroup/luisacompute` library uses a sophisticated abstraction layer to support multiple GPU backends without cluttering the codebase with conditional compilation. At the heart of this design lies the `DeviceInterface` abstraction, which establishes a strict contract that every backend must fulfill, enabling seamless platform-specific implementations behind a common interface.

## Understanding the DeviceInterface Contract

The core abstraction resides in [`include/luisa/runtime/rhi/device_interface.h`](https://github.com/luisagroup/luisacompute/blob/main/include/luisa/runtime/rhi/device_interface.h). The `DeviceInterface` class defines a pure-virtual contract covering every operation a backend must support, including resource creation, memory management, and execution control.

```cpp
// include/luisa/runtime/rhi/device_interface.h
class LUISA_RUNTIME_API DeviceInterface : public luisa::enable_shared_from_this<DeviceInterface> {
public:
    virtual ~DeviceInterface() noexcept = default;
    [[nodiscard]] virtual void *native_handle() const noexcept = 0;
    [[nodiscard]] virtual uint compute_warp_size() const noexcept = 0;
    [[nodiscard]] virtual uint64_t memory_granularity() const noexcept = 0;
    
    // Resource creation APIs
    [[nodiscard]] virtual BufferCreationInfo create_buffer(const Type *element,
                                                          size_t elem_count,
                                                          void *external_memory) noexcept = 0;
    // Stream, swap-chain, shader, and acceleration structure APIs follow...
};

```

This interface isolates the rest of the library from platform-specific details. Whether the underlying hardware uses CUDA, Metal, or DirectX, the runtime interacts solely with these virtual methods.

## How Concrete Backends Implement the Abstraction

Each supported platform implements a concrete subclass that overrides the abstract methods with calls to the native API of that platform.

### CUDA Backend Implementation

The CUDA backend implements `DeviceInterface` through the `CUDADevice` class defined in [`src/backends/cuda/cuda_device.h`](https://github.com/luisagroup/luisacompute/blob/main/src/backends/cuda/cuda_device.h). This class manages CUDA-specific contexts, compiles PTX code, and implements resource creation using CUDA driver APIs.

```cpp
// src/backends/cuda/cuda_device.h (excerpt)
class CUDADevice final : public DeviceInterface {
public:
    void *native_handle() const noexcept override { 
        return _handle.context(); 
    }
    uint compute_warp_size() const noexcept override { 
        return 32u; 
    }
    BufferCreationInfo create_buffer(const Type *element,
                                    size_t elem_count,
                                    void *external_memory) noexcept override;
    // Additional CUDA-specific implementations...
};

```

The `create_buffer` implementation allocates actual CUDA memory using `cuMemAlloc` or similar driver calls, then populates the `BufferCreationInfo` structure with device-specific handles.

### Metal and DirectX Implementations

Following the same pattern, the Metal backend provides `MetalDevice` in [`src/backends/metal/metal_device.h`](https://github.com/luisagroup/luisacompute/blob/main/src/backends/metal/metal_device.h), while the DirectX backend implements `DXDevice`. Each provides platform-specific answers to warp size queries, memory granularity, and resource creation using their respective native APIs (Metal Performance Shaders for Metal, DirectX 12 for DX).

## Runtime Device Creation and Dynamic Loading

The user-facing entry point `Context::create_device` handles backend discovery and instantiation without exposing implementation details. Defined in [`include/luisa/runtime/context.h`](https://github.com/luisagroup/luisacompute/blob/main/include/luisa/runtime/context.h) and implemented in [`src/runtime/context.cpp`](https://github.com/luisagroup/luisacompute/blob/main/src/runtime/context.cpp), this method dynamically loads the appropriate backend module.

```cpp
// include/luisa/runtime/context.h (excerpt)
[[nodiscard]] Device create_device(luisa::string_view backend_name,
                                   const DeviceConfig *settings,
                                   bool enable_validation) noexcept;

// src/runtime/context.cpp (excerpt)
auto &&m = _impl->load_backend(backend_name);          // loads shared library
Device::Creator creator = m.creator;                   // symbol: create_device(...)
auto *impl = creator(std::move(ctx), config);         // backend creates its DeviceInterface
return Device{luisa::shared_ptr<DeviceInterface>{impl}}; // wrapped in high-level Device

```

The `Device` class acts as a thin handle holding a `shared_ptr<DeviceInterface>`, forwarding all operations to the virtual interface. This design ensures that high-level code remains agnostic to whether it is running on CUDA, Metal, or any other supported platform.

## Platform-Agnostic Usage in Higher-Level Code

All higher-level modules interact with the abstraction rather than concrete implementations. For example, [`runtime/stream.cpp`](https://github.com/luisagroup/luisacompute/blob/main/runtime/stream.cpp) synchronizes execution by calling the abstract interface:

```cpp
// runtime/stream.cpp (excerpt)
void Stream::synchronize() noexcept {
    _device->synchronize_stream(_handle);
}

```

Here, `_device` is a pointer to `DeviceInterface`. The actual synchronization implementation might invoke `cuStreamSynchronize` for CUDA, `MTLCommandBuffer waitUntilCompleted` for Metal, or the equivalent DirectX command queue fence operations. The calling code requires no knowledge of these details.

## Extending Luisa Compute with New Backends

Adding support for a new GPU platform requires implementing the `DeviceInterface` contract without modifying existing runtime code. The process involves three steps:

1. **Implement the interface**: Create a subclass of `DeviceInterface` that maps every virtual function to the new platform's native API.
2. **Provide the factory**: Export a `create_device` function with the exact signature expected by the runtime loader.
3. **Register the backend**: Add the source folder to the build system (xmake) and include the backend name in the compiled-backend registry.

For example, a hypothetical Vulkan backend would look like:

```cpp
// my_vulkan_device.h
class VulkanDevice final : public DeviceInterface {
public:
    void *native_handle() const noexcept override { return _vk_device; }
    uint compute_warp_size() const noexcept override { return 32; }
    BufferCreationInfo create_buffer(const Type *element,
                                    size_t elem_count,
                                    void *external_memory) noexcept override;
    // ... remaining interface methods ...
};

// my_vulkan_device.cpp
extern "C" luisa::compute::DeviceInterface *
create_device(luisa::compute::Context &&ctx,
              const luisa::compute::DeviceConfig *cfg) noexcept {
    return new VulkanDevice(std::move(ctx), cfg);
}

```

After compilation, users can instantiate the new backend via `Context::create_device("vulkan")` without any changes to the core runtime or their application code.

## Summary

- **Pure-Virtual Contract**: `DeviceInterface` in [`include/luisa/runtime/rhi/device_interface.h`](https://github.com/luisagroup/luisacompute/blob/main/include/luisa/runtime/rhi/device_interface.h) defines the complete set of operations that any GPU backend must implement, including resource creation, memory management, and execution control.
- **Concrete Implementations**: Backends like `CUDADevice`, `MetalDevice`, and `DXDevice` inherit from `DeviceInterface` and override virtual methods with platform-specific native API calls.
- **Dynamic Loading**: The `Context::create_device` method loads backend shared libraries at runtime and instantiates concrete implementations through a factory function, wrapping them in a high-level `Device` handle.
- **Platform Agnostic**: Higher-level modules interact solely with the `DeviceInterface` abstraction, eliminating the need for conditional compilation or platform-specific code in the core library.
- **Extensibility**: New backends require only the implementation of the `DeviceInterface` contract and registration with the build system, without modifications to existing runtime code.

## Frequently Asked Questions

### What is the DeviceInterface abstraction in Luisa Compute?

The `DeviceInterface` is a pure-virtual base class defined in [`include/luisa/runtime/rhi/device_interface.h`](https://github.com/luisagroup/luisacompute/blob/main/include/luisa/runtime/rhi/device_interface.h) that serves as the runtime-level contract between the Luisa Compute library and GPU-specific backends. It declares abstract methods for every GPU operation—including buffer allocation, shader compilation, and stream synchronization—that concrete backend classes must implement using their respective native APIs (CUDA, Metal, DirectX, etc.).

### How does Luisa Compute load different GPU backends at runtime?

Luisa Compute uses dynamic loading through the `Context` class. When you call `Context::create_device("cuda")` (or "metal", "dx", etc.), the runtime loads the corresponding shared library, retrieves a factory function symbol named `create_device`, and invokes it to instantiate the backend's concrete `DeviceInterface` subclass. This factory returns a raw pointer that the `Context` wraps in a `shared_ptr` inside a high-level `Device` handle, ensuring automatic resource management while maintaining polymorphic behavior.

### Can I implement a custom backend for Luisa Compute?

Yes, you can extend Luisa Compute to support new GPU platforms or software rasterizers by implementing the `DeviceInterface` contract. You must create a subclass that overrides all pure-virtual methods with calls to your target platform's native API, then export a C-linkage factory function `create_device` that instantiates your class. After adding your source files to the xmake build system and registering the backend name, users can select your implementation via `Context::create_device("your_backend")` without any changes to the core library or their application code.

### What GPU APIs are currently supported by the DeviceInterface abstraction?

As of the current implementation, the `DeviceInterface` abstraction supports CUDA through the `CUDADevice` class in `src/backends/cuda/`, Metal through `MetalDevice` in `src/backends/metal/`, DirectX through `DXDevice` in the DirectX backend folder, and a software fallback through `FallbackDevice` in `src/backends/fallback/`. Each implementation provides platform-specific answers to queries like warp size and memory granularity while exposing unified resource creation methods to the runtime.