ncnn Memory Allocation Strategies: Understanding blob_allocator vs workspace_allocator in the Option Class
In ncnn, blob_allocator manages persistent memory for layer outputs that survive across forward passes, while workspace_allocator provides temporary scratch buffers used only during single-layer computation and immediately recycled.
The Tencent/ncnn inference framework separates memory management into two distinct strategies through the Option class. Understanding the difference between blob_allocator and workspace_allocator is essential for optimizing memory usage in deep learning deployments, particularly when implementing custom allocators or deploying on memory-constrained edge devices.
Core Differences Between blob_allocator and workspace_allocator
Purpose and Memory Lifetime
The fundamental distinction lies in data persistence:
blob_allocator: Allocates persistent storage for layer outputs (blobs). These tensors survive across layer boundaries and multiple forward passes until explicitly released when the network is cleared or the next forward pass begins.workspace_allocator: Provides temporary scratch space for intermediate calculations within a single layer. Examples includeim2coltransformations, Winograd tiles, or temporary GEMM buffers. This memory is allocated, used, and freed within the same forward call.
Default PoolAllocator Behavior
When opt.use_local_pool_allocator is true (the default) and no custom allocator is set, ncnn lazily creates separate PoolAllocator instances in Net::load_model within src/net.cpp (lines 1725-1732). The framework maintains d->local_blob_allocator for persistent data and d->local_workspace_allocator for temporary buffers, ensuring memory reuse without fragmentation between inference passes.
Implementation in Layer Code
Allocating Output Blobs
Layers allocate persistent output tensors using blob_allocator via Mat::create. In src/layer/x86/slice_x86.cpp at line 78:
top_blob.create(out_w, out_h, out_c, opt.blob_allocator);
This allocates memory that persists beyond the slice_x86 layer's execution, holding the sliced output for subsequent layers.
Allocating Workspace Buffers
For temporary computation buffers, layers use workspace_allocator. In src/layer/x86/convolution_im2col_gemm.h at line 4615:
Mat BT(..., opt.workspace_allocator);
This buffer exists only during the convolution operation. When the Mat goes out of scope at the end of the forward function, the memory returns to the workspace pool for immediate reuse by the next layer.
Configuring Custom Allocators
Implementing a Custom Allocator
Create a class inheriting from ncnn::Allocator and override fastMalloc and fastFree as defined in src/allocator.h:
class MyAllocator : public ncnn::Allocator {
public:
void* fastMalloc(size_t size) override {
// Custom allocation logic (e.g., DMA, pinned memory)
return malloc(size);
}
void fastFree(void* ptr) override {
// Custom deallocation logic
free(ptr);
}
};
Setting Allocators on Net and Extractor
Configure allocators before loading the model to control memory placement:
ncnn::Net net;
MyAllocator blob_alloc;
MyAllocator workspace_alloc;
net.opt.blob_allocator = &blob_alloc;
net.opt.workspace_allocator = &workspace_alloc;
net.opt.use_local_pool_allocator = false; // Disable default pool
net.load_model("model.bin");
Alternatively, set per-extractor at runtime for specific inference contexts:
auto ex = net.create_extractor();
ex.set_blob_allocator(&blob_alloc);
ex.set_workspace_allocator(&workspace_alloc);
Vulkan GPU Allocators
For GPU inference, use the Vulkan-specific variants declared in src/option.h (lines 48 and 51):
ncnn::VulkanDevice* vkdev = ncnn::get_gpu_device(0);
ncnn::VkBlobAllocator vk_blob_alloc(vkdev);
ncnn::VkWeightAllocator vk_workspace_alloc(vkdev); // Can serve as workspace
net.opt.blob_vkallocator = &vk_blob_alloc;
net.opt.workspace_vkallocator = &vk_workspace_alloc;
Summary
blob_allocatormanages persistent memory for layer outputs that survive across forward passes and network boundaries, allocated viaMat::createin layer implementations.workspace_allocatorprovides temporary scratch buffers used only during single-layer computation (e.g.,im2col, Winograd transforms) and immediately recycled.- Both default to separate
PoolAllocatorinstances whenuse_local_pool_allocatoris enabled, created lazily insrc/net.cppduring model loading. - Custom allocators can be injected via
ncnn::Optionor per-extractor methods for specialized memory management such as DMA, NUMA-aware allocation, or memory tracing.
Frequently Asked Questions
Can I use the same allocator instance for both blob and workspace memory?
While technically possible, it is not recommended. The blob_allocator and workspace_allocator have fundamentally different lifetime patterns—persistent versus temporary. Using separate allocators allows the framework to optimize memory reuse and prevents fragmentation that would occur if temporary buffers competed with persistent tensors in the same pool. According to the ncnn source code in src/net.cpp, the framework explicitly maintains separate pool allocators for each type when use_local_pool_allocator is true.
How does use_local_pool_allocator affect performance?
When use_local_pool_allocator is true (the default), ncnn creates dedicated PoolAllocator instances for both allocator types in src/net.cpp (lines 1725-1732). These pool allocators cache memory blocks between forward passes, reducing system allocation overhead and improving cache locality. Disabling this forces every allocation to hit the system allocator via fastMalloc/fastFree, which can significantly degrade performance in latency-sensitive applications or when processing variable input sizes that trigger frequent reallocations.
Are Vulkan allocators compatible with CPU allocators?
No, VkBlobAllocator and VkWeightAllocator manage GPU device memory and are distinct from CPU allocators. They are set via blob_vkallocator and workspace_vkallocator in src/option.h (lines 48 and 51), whereas CPU allocators use blob_allocator and workspace_allocator. You cannot assign a CPU allocator to a GPU slot or vice versa without type errors, as they inherit from different base classes (VkAllocator vs Allocator) and handle fundamentally different memory spaces (device vs host).
When should I implement a custom allocator?
Implement a custom allocator when you need specialized memory management such as memory mapping, DMA buffer allocation, NUMA-aware allocation, or memory tracing. For example, embedded systems might require allocation from specific physical memory regions, or profiling tools might need to track allocation patterns to detect leaks. Inherit from ncnn::Allocator (defined in src/allocator.h) and override fastMalloc and fastFree, then inject your implementation via net.opt or per-extractor configuration before inference begins.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →