How ncnn Manages GPU Memory with Vulkan: Architecture and Performance Trade-offs vs CPU Inference

ncnn abstracts Vulkan GPU memory through a hierarchy of allocator classes that pool device memory in blocks, support zero-copy on unified memory architectures, and minimize allocation overhead—delivering 5-30× faster compute throughput than CPU inference at the cost of higher memory footprint and initial setup latency.

Tencent's ncnn is a high-performance neural network inference framework optimized for mobile and edge devices. When leveraging Vulkan for GPU acceleration, ncnn implements a sophisticated memory management system that balances allocation efficiency with hardware-specific optimizations. This article examines how ncnn manages GPU memory with Vulkan and analyzes the performance trade-offs when comparing GPU inference against traditional CPU execution.

Vulkan Memory Architecture in ncnn

The VkAllocator Interface

At the core of ncnn's Vulkan memory management is the abstract VkAllocator class defined in src/allocator.h. This interface provides the contract between the inference engine and Vulkan device memory, exposing methods for buffer and image creation while hiding driver-specific complexity. Each allocator maintains a reference to a VulkanDevice* and implements virtual methods for memory acquisition and release.

Low-Level Allocation Primitives

The base implementation in src/allocator.cpp provides four critical allocation primitives that concrete allocators leverage:

  • allocate_memory (allocator.cpp#L46-L63) – Wraps vkAllocateMemory for generic memory blocks using standard Vulkan memory types.
  • allocate_dedicated_memory (allocator.cpp#L65-L88) – Uses VkMemoryDedicatedAllocateInfoKHR to bind individual images or buffers to dedicated allocations, required for large resources and certain discrete GPU implementations.
  • allocate_import_host_memory (allocator.cpp#L91-L115) – Imports host-accessible memory via VK_EXTERNAL_MEMORY_HANDLE_TYPE_HOST_ALLOCATION_BIT_EXT, enabling CPU writes directly into GPU-visible memory without staging buffers.
  • create_buffer / create_image / create_imageview – Thin wrappers around Vulkan object creation that apply device-specific alignment requirements.

High-Level Allocator Implementations

ncnn provides specialized allocator classes that implement pooling and caching strategies atop the low-level primitives:

VkBlobAllocator pools device-local memory in large blocks (default 16 MiB) and sub-allocates tensors from these blocks. It maintains per-queue budgets and aligns allocations to the device's buffer_offset_alignment and buffer_image_granularity. The constructor at allocator.cpp#L110-L124 calculates least-common-multiple alignments for integrated GPUs with multiple constraint requirements.

VkWeightAllocator optimizes for static model weights, optionally preferring host-visible memory when the device exposes coherent host memory, eliminating copy steps for immutable parameters.

VkStagingAllocator manages host-visible staging buffers for upload/download operations when compute memory is device-only. It supports dynamic sizing via set_size_compare_ratio (default 0.75) to balance memory usage versus allocation frequency.

VkAndroidHardwareBufferImageAllocator provides platform-specific support for importing Android AHardwareBuffer objects directly as Vulkan images, avoiding CPU-side pixel buffers on Android devices.

Zero-Copy and Unified Memory Optimization

When running on unified memory architectures (typically integrated GPUs), ncnn leverages allocate_import_host_memory to eliminate data transfers between host and device. In this mode, the application can map a VkMat directly using the mapped() method:

// Zero-copy memory access on unified memory GPUs
ncnn::VkMat blob_gpu(224, 224, 3, 4u, vkdev->acquire_blob_allocator());

// Direct CPU write into GPU memory - no staging buffer needed
ncnn::Mat host_view = blob_gpu.mapped();
memcpy(host_view.data, raw_pixels, host_view.total() * host_view.elemsize);

// Use blob_gpu directly in inference - zero copy overhead
ex.input("data", blob_gpu);

This approach removes explicit flush and invalidate operations and eliminates the extra memcpy typically required when using discrete GPUs with separate device-local memory.

Memory Lifecycle and Reclamation

ncnn optimizes allocator lifecycle management through the VulkanDevice class in src/gpu.cpp. The device maintains pools of allocators per queue family, allowing reuse across inference sessions:

Acquire / Reclaim Pattern:

// Acquire allocators from the device pool
ncnn::VkAllocator* blob_vkallocator = vkdev->acquire_blob_allocator();
ncnn::VkAllocator* staging_vkallocator = vkdev->acquire_staging_allocator();

// Configure network options
net.opt.blob_vkallocator = blob_vkallocator;
net.opt.workspace_vkallocator = blob_vkallocator;
net.opt.staging_vkallocator = staging_vkallocator;

// ... run inference ...

// Return allocators to pool for reuse
vkdev->reclaim_blob_allocator(blob_vkallocator);
vkdev->reclaim_staging_allocator(staging_vkallocator);

This pattern avoids repeated new/delete operations and Vulkan allocation calls, which are relatively expensive compared to standard memory allocation.

Dummy Object Warm-up: At device initialization, ncnn allocates a minimal dummy buffer and image (create_dummy_buffer_image() in gpu.cpp#L32-L55) to "warm up" the Vulkan driver. This ensures that subsequent command buffer submissions during actual inference avoid first-time initialization overhead.

Performance Trade-offs: GPU vs CPU Inference

When deciding between Vulkan GPU and CPU inference in ncnn, consider the following architectural differences:

Compute Throughput GPU inference leverages massive parallelism for matrix operations (convolution, GEMM), typically achieving 5-30× speedup over CPU on modern discrete and integrated GPUs. CPU inference is limited to core count and SIMD width, suitable only for very small workloads or low-batch scenarios.

Memory Bandwidth Device-local GPU memory provides 200-500 GB/s bandwidth versus ~50 GB/s for system RAM. However, CPU inference avoids device-to-host copy penalties entirely, as data remains in system memory throughout execution.

Allocation Overhead Vulkan memory allocations (vkAllocateMemory) are heavyweight operations. ncnn mitigates this through VkBlobAllocator block pooling, but GPU inference still incurs higher initial setup costs compared to CPU malloc/new operations, which are cheap and reuse std::vector buffers without driver calls.

Data Transfer Costs Unless running on unified memory architectures, GPU inference requires host-to-device copies. ncnn hides this complexity with VkStagingAllocator, but an extra memcpy remains unavoidable on discrete GPUs. CPU inference requires zero data movement.

Latency Characteristics GPU inference adds 0.5-2 ms overhead from command buffer submission and driver synchronization, making it unsuitable for micro-batches requiring sub-millisecond latency. CPU inference provides immediate execution with lower latency for tiny inputs.

Memory Footprint GPU inference often doubles memory usage by holding both device-local copies and optional staging buffers. CPU inference maintains only a single copy in RAM, yielding smaller footprint.

Power Consumption GPUs may draw more power during active computation, but finish faster (energy-time trade-off). CPUs generally exhibit lower peak power but require longer runtimes for equivalent work.

Practical Implementation Examples

Configuring Custom Allocators for GPU Inference

#include <net.h>

// Initialize Vulkan device (first available GPU)
ncnn::VulkanDevice* vkdev = ncnn::get_default_vkdev();

// Acquire allocators from device pools
ncnn::VkAllocator* blob_alloc = vkdev->acquire_blob_allocator();
ncnn::VkAllocator* staging_alloc = vkdev->acquire_staging_allocator();

// Configure network for Vulkan compute
ncnn::Net net;
net.opt.use_vulkan_compute = true;
net.opt.blob_vkallocator = blob_alloc;
net.opt.workspace_vkallocator = blob_alloc;
net.opt.staging_vkallocator = staging_alloc;

// Load model weights (must occur after Vulkan configuration)
net.load_param("mobilenet_v2.param");
net.load_model("mobilenet_v2.bin");

// Create extractor and run inference
ncnn::Extractor ex = net.create_extractor();
ex.input("data", ncnn::Mat::from_pixels_resize(image.data, ncnn::Mat::PIXEL_RGB, 224, 224));

ncnn::Mat out;
ex.extract("prob", out);

// Return allocators to pool for reuse
vkdev->reclaim_blob_allocator(blob_alloc);
vkdev->reclaim_staging_allocator(staging_alloc);

Zero-Copy Memory on Unified Architectures

// Check if device supports unified memory (integrated GPUs)
if (vkdev->info.type == ncnn::VulkanDevice::DeviceType::INTEGRATED_GPU) {
    // Create VkMat with blob allocator
    ncnn::VkMat blob_gpu(224, 224, 3, 4u, vkdev->acquire_blob_allocator());
    
    // Map directly into host-accessible GPU memory
    ncnn::Mat host_view = blob_gpu.mapped();
    
    // Write pixel data directly (no staging buffer, no memcpy)
    memcpy(host_view.data, raw_pixel_data, host_view.total() * host_view.elemsize);
    
    // Use directly in inference - zero copy overhead
    ex.input("data", blob_gpu);
}

Hybrid CPU-GPU Pipeline

// Setup CPU network for preprocessing
ncnn::Net net_cpu;
net_cpu.opt.use_vulkan_compute = false;
net_cpu.load_param("preproc.param");
net_cpu.load_model("preproc.bin");

// Setup GPU network for main inference
ncnn::Net net_gpu;
net_gpu.opt.use_vulkan_compute = true;
net_gpu.opt.blob_vkallocator = vkdev->acquire_blob_allocator();
net_gpu.opt.staging_vkallocator = vkdev->acquire_staging_allocator();
net_gpu.load_param("model.param");
net_gpu.load_model("model.bin");

// Step 1: CPU preprocessing
ncnn::Extractor ex_cpu = net_cpu.create_extractor();
ex_cpu.input("input", raw_image);
ncnn::Mat preprocessed;
ex_cpu.extract("output", preprocessed);

// Step 2: Upload to GPU using staging allocator
ncnn::VkMat gpu_input;
gpu_input.create_like(preprocessed, net_gpu.opt.blob_vkallocator);
gpu_input.upload(preprocessed, net_gpu.opt.staging_vkallocator);

// Step 3: GPU inference
ncnn::Extractor ex_gpu = net_gpu.create_extractor();
ex_gpu.input("data", gpu_input);
ncnn::VkMat gpu_output;
ex_gpu.extract("prob", gpu_output);

// Cleanup
vkdev->reclaim_blob_allocator(net_gpu.opt.blob_vkallocator);
vkdev->reclaim_staging_allocator(net_gpu.opt.staging_vkallocator);

Summary

  • ncnn manages Vulkan GPU memory through a hierarchy of allocator classes centered on the VkAllocator interface, with concrete implementations like VkBlobAllocator, VkWeightAllocator, and VkStagingAllocator providing block pooling, dedicated allocations, and host-imported memory support.
  • Block-wise pooling minimizes driver overhead by sub-allocating from large 16 MiB chunks rather than calling vkAllocateMemory for every tensor, significantly reducing allocation latency during inference.
  • Zero-copy optimization via allocate_import_host_memory eliminates staging buffers on unified memory architectures (integrated GPUs), allowing direct CPU writes into GPU-visible memory through the mapped() method.
  • GPU inference offers 5-30× compute throughput for matrix operations but incurs higher memory footprint, 0.5-2 ms submission latency, and requires careful management of staging buffers on discrete GPUs.
  • CPU inference remains preferable for micro-batches requiring sub-millisecond latency, memory-constrained environments, or when running on devices without Vulkan compute support.

Frequently Asked Questions

What is the primary benefit of using Vulkan for GPU inference in ncnn?

The primary benefit is massive parallel compute throughput for deep learning operations like convolution and GEMM. According to the ncnn source code, Vulkan GPU inference can achieve 5-30× speedup over CPU inference on modern hardware by leveraging hundreds of compute cores simultaneously. The VkBlobAllocator and specialized shader pipelines in src/layer/vulkan/shader/ ensure that tensor data remains in high-bandwidth device-local memory (~200-500 GB/s) throughout the computation.

How does ncnn handle memory allocation on discrete versus integrated GPUs?

ncnn adapts its allocation strategy based on the device's memory architecture through the allocate_import_host_memory primitive. On integrated GPUs with unified memory, ncnn imports host-accessible memory via VK_EXTERNAL_MEMORY_HANDLE_TYPE_HOST_ALLOCATION_BIT_EXT, enabling zero-copy access where CPU writes go directly to GPU memory. On discrete GPUs with separate device-local memory, ncnn uses VkStagingAllocator to manage host-visible staging buffers for uploads/downloads, while VkBlobAllocator handles device-local memory pooling. The VkWeightAllocator can optionally keep model weights in host-visible memory on discrete GPUs to avoid redundant copies if the memory is coherent.

When should I choose CPU inference over GPU inference in ncnn?

Choose CPU inference when your workload requires sub-millisecond latency for tiny batches, runs on memory-constrained devices where GPU memory doubling would exceed available RAM, or targets hardware without Vulkan compute support. The ncnn source shows that GPU inference adds 0.5-2 ms of overhead from command buffer submission and driver synchronization, making it unsuitable for micro-batches. Additionally, if your model is small and the CPU can process it within thermal and power constraints without waking the GPU (which may draw significant power), CPU inference provides better energy efficiency for sporadic inference tasks.

Can ncnn use zero-copy memory on all Vulkan-capable devices?

No, zero-copy memory via allocate_import_host_memory is only available on devices that expose a unified memory architecture and support the VK_EXTERNAL_MEMORY_HANDLE_TYPE_HOST_ALLOCATION_BIT_EXT extension. This typically includes integrated GPUs like Intel UHD Graphics or Apple Silicon, where CPU and GPU share physical memory. On discrete GPUs (NVIDIA GeForce, AMD Radeon), memory is physically separate, requiring explicit staging buffers managed by VkStagingAllocator. You can check for zero-copy support by attempting to map a VkMat using the mapped() method—if the device doesn't support host-imported memory, ncnn will fall back to traditional staging buffer uploads.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →