# #LuisaCompute GPU Memory Management Strategy: Vulkan, Metal, CUDA, and CPU Implementations

> Explore LuisaCompute's GPU memory management strategy across Vulkan, Metal, CUDA, and CPU. Discover unified resource abstraction and platform-specific allocation layers.

- Repository: [LuisaGroup/luisacompute](https://github.com/luisagroup/luisacompute)
- Tags: internals
- Published: 2026-03-06

---

#LuisaCompute GPU Memory Management Strategy: Vulkan, Metal, CUDA, and CPU Implementations

**LuisaCompute abstracts GPU resources behind a unified Backend interface, with each backend implementing platform-specific allocation layers—VMA for Vulkan/CUDA interop, MTLHeap pools for Metal, cudaMalloc for native CUDA, and host allocators for CPU validation—while exposing a common API.**

LuisaCompute is a high-performance computing framework developed by **luisagroup/luisacompute** that unifies GPU programming across heterogeneous hardware. Understanding its **GPU memory management strategy** requires examining how each backend adapts native APIs—whether Vulkan Memory Allocator for cross-platform interop or Metal Heaps for Apple silicon—into a coherent device interface.

## Vulkan Backend: VMA-Based Sub-Allocation

The Vulkan backend (servicing CUDA-Vulkan interop scenarios) leverages **Vulkan Memory Allocator (VMA)** to manage `VkDeviceMemory` blocks efficiently. This approach minimizes driver overhead by sub-allocating resources within larger memory blocks rather than requesting individual allocations from the driver.

### Allocation Workflow and Flags

In [`src/backends/vk/vk_mem_alloc.h`](https://github.com/luisagroup/luisacompute/blob/main/src/backends/vk/vk_mem_alloc.h), the backend instantiates a single `VmaAllocator` during device initialization. When creating resources, LuisaCompute invokes `vmaCreateBuffer` or `vmaCreateImage` (implemented in [`src/backends/vk/buffer.cpp`](https://github.com/luisagroup/luisacompute/blob/main/src/backends/vk/buffer.cpp)), passing strategy-specific flags that control allocation behavior.

Key configuration flags include:
- `VMA_ALLOCATION_CREATE_DEDICATED_MEMORY_BIT` for large resources requiring isolated memory blocks
- `VMA_ALLOCATION_CREATE_LINEAR_ALGORITHM_BIT` for linear sub-allocation pools suitable for transient buffers
- Support for `VK_KHR_dedicated_allocation` and `VK_EXT_memory_priority` extensions to optimize memory placement

### Dedicated vs. Linear Allocation Strategies

VMA selects between strategies based on resource size and usage patterns. Small buffers are sub-allocated from existing `VkDeviceMemory` blocks using best-fit or linear algorithms, while large textures or buffers flagged for dedication receive their own memory blocks. The library exposes allocation statistics (used bytes, fragmentation ratios) which LuisaCompute logs for performance profiling.

## Metal Backend: Stage-Buffer Pool Architecture

Apple platforms utilize **Metal Heaps** and a specialized staging buffer pool to minimize allocation overhead and reduce CPU-GPU synchronization points.

### MTLHeap Pre-Allocation Configuration

According to [`src/backends/metal/metal_stage_buffer_pool.cpp`](https://github.com/luisagroup/luisacompute/blob/main/src/backends/metal/metal_stage_buffer_pool.cpp), the backend creates a large `MTLHeap` at device startup, sized by the `--lc-metal-stage-buffer-pool-size` configuration option. This pool dishes out sub-buffers via `newBuffer` calls for transient staging data (uploads/downloads), avoiding expensive heap expansion during frame rendering.

### Storage Modes for Buffers and Textures

Buffers allocate with `MTLResourceStorageModeShared` for CPU-GPU coherent access or `MTLResourceStorageModePrivate` when data is GPU-read-only. For textures, the backend calls `newTexture` directly in [`src/backends/metal/metal_device.cpp`](https://github.com/luisagroup/luisacompute/blob/main/src/backends/metal/metal_device.cpp), allocating large resources as private to eliminate unnecessary CPU-GPU synchronization overhead.

## CUDA Backend: Native Allocation and External Memory Interop

The CUDA backend provides both direct device allocation for native execution and external memory sharing mechanisms for cross-backend workflows.

### Runtime Allocation Wrappers

Standard allocations use `cudaMalloc` for device-only memory or `cudaMallocManaged` for unified memory accessible from both host and device. These are thinly wrapped in headers like [`src/backends/cuda/cuda_allocator.h`](https://github.com/luisagroup/luisacompute/blob/main/src/backends/cuda/cuda_allocator.h), maintaining alignment and stream-ordered deallocation semantics.

### Vulkan Interop Mechanisms

For memory sharing with Vulkan, [`src/backends/cuda/cuda_external_memory.cpp`](https://github.com/luisagroup/luisacompute/blob/main/src/backends/cuda/cuda_external_memory.cpp) implements `cudaImportExternalMemory` to import Vulkan `VkDeviceMemory` handles exported via `VK_KHR_external_memory`. The function `cudaExternalMemoryGetMappedBuffer` returns CUDA device pointers that map to the same physical memory backing Vulkan resources, enabling zero-copy data exchange between compute and graphics pipelines.

## CPU Validation Backend: Host Memory Abstraction

The validation backend enables kernel testing without GPU hardware. In [`src/backends/validation/host_allocator.h`](https://github.com/luisagroup/luisacompute/blob/main/src/backends/validation/host_allocator.h), it implements the `Device` interface using standard host allocators (`malloc` or `std::vector<std::byte>`). Because memory is host-visible, mapping operations are no-ops, and kernels execute through a software rasterizer that interprets the kernel intermediate representation.

## Cross-Backend Usage Examples

Despite differing allocation strategies, the public API remains consistent across all backends:

```cpp
// Unified API works across all backends
luisa::compute::Buffer<float> buf = device.create_buffer<float>(1024);

// Vulkan: VMA sub-allocation or dedicated path
auto *vk_dev = static_cast<luisa::compute::vk::VulkanDevice*>(device.native_handle());
auto vk_buf = vk_dev->allocator()->create_buffer(
    sizeof(float) * 1024,
    VMA_ALLOCATION_CREATE_DEDICATED_MEMORY_BIT);

// Metal: Stage-buffer pool allocation
auto *mtl_dev = static_cast<luisa::compute::metal::MetalDevice*>(device.native_handle());
auto mtl_buf = mtl_dev->stage_buffer_pool()->allocate(sizeof(float) * 1024);

// CUDA: Direct runtime allocation
float *cuda_ptr;
cudaMalloc(&cuda_ptr, sizeof(float) * 1024);

```

## Summary

- **Vulkan Backend**: Uses the VMA library in [`src/backends/vk/vk_mem_alloc.h`](https://github.com/luisagroup/luisacompute/blob/main/src/backends/vk/vk_mem_alloc.h) for sub-allocation within `VkDeviceMemory` blocks, supporting dedicated allocations for large resources and linear pools for transient data.
- **Metal Backend**: Implements heap-based pooling via `MTLHeap` in [`src/backends/metal/metal_stage_buffer_pool.cpp`](https://github.com/luisagroup/luisacompute/blob/main/src/backends/metal/metal_stage_buffer_pool.cpp), configurable via `--lc-metal-stage-buffer-pool-size`, with separate storage modes for staging and private GPU data.
- **CUDA Backend**: Wraps `cudaMalloc` and `cudaMallocManaged` for native execution, while [`src/backends/cuda/cuda_external_memory.cpp`](https://github.com/luisagroup/luisacompute/blob/main/src/backends/cuda/cuda_external_memory.cpp) handles Vulkan interop through `cudaImportExternalMemory` and mapped buffer handles.
- **CPU Backend**: Provides host memory allocation in [`src/backends/validation/host_allocator.h`](https://github.com/luisagroup/luisacompute/blob/main/src/backends/validation/host_allocator.h) for software rasterization and unit testing without physical GPU hardware.

## Frequently Asked Questions

### How does LuisaCompute minimize memory fragmentation on Vulkan?

The Vulkan backend relies on VMA's built-in tracking statistics to monitor fragmentation levels across `VkDeviceMemory` blocks. It mitigates fragmentation by selecting between dedicated allocations (via `VMA_ALLOCATION_CREATE_DEDICATED_MEMORY_BIT`) for large persistent resources and linear sub-allocation pools (via `VMA_ALLOCATION_CREATE_LINEAR_ALGORITHM_BIT`) for smaller transient buffers, as implemented in [`src/backends/vk/vk_mem_alloc.h`](https://github.com/luisagroup/luisacompute/blob/main/src/backends/vk/vk_mem_alloc.h).

### Can Metal buffers be shared directly with CUDA?

No, direct Metal-to-CUDA memory sharing is not implemented in the current stable branch. LuisaCompute uses Vulkan as the interoperability bridge: CUDA shares memory with Vulkan through external memory handles imported via `cudaImportExternalMemory` in [`src/backends/cuda/cuda_external_memory.cpp`](https://github.com/luisagroup/luisacompute/blob/main/src/backends/cuda/cuda_external_memory.cpp), while Metal operates independently via its `MTLHeap` pool system.

### What is the purpose of the CPU validation backend?

The CPU backend exists primarily for unit testing and kernel debugging without requiring physical GPU hardware. As implemented in [`src/backends/validation/host_allocator.h`](https://github.com/luisagroup/luisacompute/blob/main/src/backends/validation/host_allocator.h), it allocates host memory using standard system allocators and executes kernels through a software interpreter, making buffer mapping operations zero-cost while validating kernel logic.

### How do you configure the Metal stage-buffer pool size?

Set the `--lc-metal-stage-buffer-pool-size` command-line option or configuration parameter before device initialization. This value determines the initial capacity of the `MTLHeap` created in [`src/backends/metal/metal_device.cpp`](https://github.com/luisagroup/luisacompute/blob/main/src/backends/metal/metal_device.cpp) and [`metal_stage_buffer_pool.cpp`](https://github.com/luisagroup/luisacompute/blob/main/metal_stage_buffer_pool.cpp), controlling how many transient staging buffers can be allocated before the heap requires expansion.