#LuisaCompute GPU Memory Management Strategy: Vulkan, Metal, CUDA, and CPU Implementations

#LuisaCompute GPU Memory Management Strategy: Vulkan, Metal, CUDA, and CPU Implementations

LuisaCompute abstracts GPU resources behind a unified Backend interface, with each backend implementing platform-specific allocation layers—VMA for Vulkan/CUDA interop, MTLHeap pools for Metal, cudaMalloc for native CUDA, and host allocators for CPU validation—while exposing a common API.

LuisaCompute is a high-performance computing framework developed by luisagroup/luisacompute that unifies GPU programming across heterogeneous hardware. Understanding its GPU memory management strategy requires examining how each backend adapts native APIs—whether Vulkan Memory Allocator for cross-platform interop or Metal Heaps for Apple silicon—into a coherent device interface.

Vulkan Backend: VMA-Based Sub-Allocation

The Vulkan backend (servicing CUDA-Vulkan interop scenarios) leverages Vulkan Memory Allocator (VMA) to manage VkDeviceMemory blocks efficiently. This approach minimizes driver overhead by sub-allocating resources within larger memory blocks rather than requesting individual allocations from the driver.

Allocation Workflow and Flags

In src/backends/vk/vk_mem_alloc.h, the backend instantiates a single VmaAllocator during device initialization. When creating resources, LuisaCompute invokes vmaCreateBuffer or vmaCreateImage (implemented in src/backends/vk/buffer.cpp), passing strategy-specific flags that control allocation behavior.

Key configuration flags include:

  • VMA_ALLOCATION_CREATE_DEDICATED_MEMORY_BIT for large resources requiring isolated memory blocks
  • VMA_ALLOCATION_CREATE_LINEAR_ALGORITHM_BIT for linear sub-allocation pools suitable for transient buffers
  • Support for VK_KHR_dedicated_allocation and VK_EXT_memory_priority extensions to optimize memory placement

Dedicated vs. Linear Allocation Strategies

VMA selects between strategies based on resource size and usage patterns. Small buffers are sub-allocated from existing VkDeviceMemory blocks using best-fit or linear algorithms, while large textures or buffers flagged for dedication receive their own memory blocks. The library exposes allocation statistics (used bytes, fragmentation ratios) which LuisaCompute logs for performance profiling.

Metal Backend: Stage-Buffer Pool Architecture

Apple platforms utilize Metal Heaps and a specialized staging buffer pool to minimize allocation overhead and reduce CPU-GPU synchronization points.

MTLHeap Pre-Allocation Configuration

According to src/backends/metal/metal_stage_buffer_pool.cpp, the backend creates a large MTLHeap at device startup, sized by the --lc-metal-stage-buffer-pool-size configuration option. This pool dishes out sub-buffers via newBuffer calls for transient staging data (uploads/downloads), avoiding expensive heap expansion during frame rendering.

Storage Modes for Buffers and Textures

Buffers allocate with MTLResourceStorageModeShared for CPU-GPU coherent access or MTLResourceStorageModePrivate when data is GPU-read-only. For textures, the backend calls newTexture directly in src/backends/metal/metal_device.cpp, allocating large resources as private to eliminate unnecessary CPU-GPU synchronization overhead.

CUDA Backend: Native Allocation and External Memory Interop

The CUDA backend provides both direct device allocation for native execution and external memory sharing mechanisms for cross-backend workflows.

Runtime Allocation Wrappers

Standard allocations use cudaMalloc for device-only memory or cudaMallocManaged for unified memory accessible from both host and device. These are thinly wrapped in headers like src/backends/cuda/cuda_allocator.h, maintaining alignment and stream-ordered deallocation semantics.

Vulkan Interop Mechanisms

For memory sharing with Vulkan, src/backends/cuda/cuda_external_memory.cpp implements cudaImportExternalMemory to import Vulkan VkDeviceMemory handles exported via VK_KHR_external_memory. The function cudaExternalMemoryGetMappedBuffer returns CUDA device pointers that map to the same physical memory backing Vulkan resources, enabling zero-copy data exchange between compute and graphics pipelines.

CPU Validation Backend: Host Memory Abstraction

The validation backend enables kernel testing without GPU hardware. In src/backends/validation/host_allocator.h, it implements the Device interface using standard host allocators (malloc or std::vector<std::byte>). Because memory is host-visible, mapping operations are no-ops, and kernels execute through a software rasterizer that interprets the kernel intermediate representation.

Cross-Backend Usage Examples

Despite differing allocation strategies, the public API remains consistent across all backends:

// Unified API works across all backends
luisa::compute::Buffer<float> buf = device.create_buffer<float>(1024);

// Vulkan: VMA sub-allocation or dedicated path
auto *vk_dev = static_cast<luisa::compute::vk::VulkanDevice*>(device.native_handle());
auto vk_buf = vk_dev->allocator()->create_buffer(
    sizeof(float) * 1024,
    VMA_ALLOCATION_CREATE_DEDICATED_MEMORY_BIT);

// Metal: Stage-buffer pool allocation
auto *mtl_dev = static_cast<luisa::compute::metal::MetalDevice*>(device.native_handle());
auto mtl_buf = mtl_dev->stage_buffer_pool()->allocate(sizeof(float) * 1024);

// CUDA: Direct runtime allocation
float *cuda_ptr;
cudaMalloc(&cuda_ptr, sizeof(float) * 1024);

Summary

  • Vulkan Backend: Uses the VMA library in src/backends/vk/vk_mem_alloc.h for sub-allocation within VkDeviceMemory blocks, supporting dedicated allocations for large resources and linear pools for transient data.
  • Metal Backend: Implements heap-based pooling via MTLHeap in src/backends/metal/metal_stage_buffer_pool.cpp, configurable via --lc-metal-stage-buffer-pool-size, with separate storage modes for staging and private GPU data.
  • CUDA Backend: Wraps cudaMalloc and cudaMallocManaged for native execution, while src/backends/cuda/cuda_external_memory.cpp handles Vulkan interop through cudaImportExternalMemory and mapped buffer handles.
  • CPU Backend: Provides host memory allocation in src/backends/validation/host_allocator.h for software rasterization and unit testing without physical GPU hardware.

Frequently Asked Questions

How does LuisaCompute minimize memory fragmentation on Vulkan?

The Vulkan backend relies on VMA's built-in tracking statistics to monitor fragmentation levels across VkDeviceMemory blocks. It mitigates fragmentation by selecting between dedicated allocations (via VMA_ALLOCATION_CREATE_DEDICATED_MEMORY_BIT) for large persistent resources and linear sub-allocation pools (via VMA_ALLOCATION_CREATE_LINEAR_ALGORITHM_BIT) for smaller transient buffers, as implemented in src/backends/vk/vk_mem_alloc.h.

Can Metal buffers be shared directly with CUDA?

No, direct Metal-to-CUDA memory sharing is not implemented in the current stable branch. LuisaCompute uses Vulkan as the interoperability bridge: CUDA shares memory with Vulkan through external memory handles imported via cudaImportExternalMemory in src/backends/cuda/cuda_external_memory.cpp, while Metal operates independently via its MTLHeap pool system.

What is the purpose of the CPU validation backend?

The CPU backend exists primarily for unit testing and kernel debugging without requiring physical GPU hardware. As implemented in src/backends/validation/host_allocator.h, it allocates host memory using standard system allocators and executes kernels through a software interpreter, making buffer mapping operations zero-cost while validating kernel logic.

How do you configure the Metal stage-buffer pool size?

Set the --lc-metal-stage-buffer-pool-size command-line option or configuration parameter before device initialization. This value determines the initial capacity of the MTLHeap created in src/backends/metal/metal_device.cpp and metal_stage_buffer_pool.cpp, controlling how many transient staging buffers can be allocated before the heap requires expansion.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →