ncnn Memory Allocation Strategies: Understanding blob_allocator vs workspace_allocator in the Option Class

In ncnn, blob_allocator manages persistent memory for layer outputs that survive across forward passes, while workspace_allocator provides temporary scratch buffers used only during single-layer computation and immediately recycled.

The Tencent/ncnn inference framework separates memory management into two distinct strategies through the Option class. Understanding the difference between blob_allocator and workspace_allocator is essential for optimizing memory usage in deep learning deployments, particularly when implementing custom allocators or deploying on memory-constrained edge devices.

Core Differences Between blob_allocator and workspace_allocator

Purpose and Memory Lifetime

The fundamental distinction lies in data persistence:

  • blob_allocator: Allocates persistent storage for layer outputs (blobs). These tensors survive across layer boundaries and multiple forward passes until explicitly released when the network is cleared or the next forward pass begins.
  • workspace_allocator: Provides temporary scratch space for intermediate calculations within a single layer. Examples include im2col transformations, Winograd tiles, or temporary GEMM buffers. This memory is allocated, used, and freed within the same forward call.

Default PoolAllocator Behavior

When opt.use_local_pool_allocator is true (the default) and no custom allocator is set, ncnn lazily creates separate PoolAllocator instances in Net::load_model within src/net.cpp (lines 1725-1732). The framework maintains d->local_blob_allocator for persistent data and d->local_workspace_allocator for temporary buffers, ensuring memory reuse without fragmentation between inference passes.

Implementation in Layer Code

Allocating Output Blobs

Layers allocate persistent output tensors using blob_allocator via Mat::create. In src/layer/x86/slice_x86.cpp at line 78:

top_blob.create(out_w, out_h, out_c, opt.blob_allocator);

This allocates memory that persists beyond the slice_x86 layer's execution, holding the sliced output for subsequent layers.

Allocating Workspace Buffers

For temporary computation buffers, layers use workspace_allocator. In src/layer/x86/convolution_im2col_gemm.h at line 4615:

Mat BT(..., opt.workspace_allocator);

This buffer exists only during the convolution operation. When the Mat goes out of scope at the end of the forward function, the memory returns to the workspace pool for immediate reuse by the next layer.

Configuring Custom Allocators

Implementing a Custom Allocator

Create a class inheriting from ncnn::Allocator and override fastMalloc and fastFree as defined in src/allocator.h:

class MyAllocator : public ncnn::Allocator {
public:
    void* fastMalloc(size_t size) override {
        // Custom allocation logic (e.g., DMA, pinned memory)
        return malloc(size);
    }
    void fastFree(void* ptr) override {
        // Custom deallocation logic
        free(ptr);
    }
};

Setting Allocators on Net and Extractor

Configure allocators before loading the model to control memory placement:

ncnn::Net net;
MyAllocator blob_alloc;
MyAllocator workspace_alloc;

net.opt.blob_allocator = &blob_alloc;
net.opt.workspace_allocator = &workspace_alloc;
net.opt.use_local_pool_allocator = false;  // Disable default pool

net.load_model("model.bin");

Alternatively, set per-extractor at runtime for specific inference contexts:

auto ex = net.create_extractor();
ex.set_blob_allocator(&blob_alloc);
ex.set_workspace_allocator(&workspace_alloc);

Vulkan GPU Allocators

For GPU inference, use the Vulkan-specific variants declared in src/option.h (lines 48 and 51):

ncnn::VulkanDevice* vkdev = ncnn::get_gpu_device(0);
ncnn::VkBlobAllocator vk_blob_alloc(vkdev);
ncnn::VkWeightAllocator vk_workspace_alloc(vkdev);  // Can serve as workspace

net.opt.blob_vkallocator = &vk_blob_alloc;
net.opt.workspace_vkallocator = &vk_workspace_alloc;

Summary

  • blob_allocator manages persistent memory for layer outputs that survive across forward passes and network boundaries, allocated via Mat::create in layer implementations.
  • workspace_allocator provides temporary scratch buffers used only during single-layer computation (e.g., im2col, Winograd transforms) and immediately recycled.
  • Both default to separate PoolAllocator instances when use_local_pool_allocator is enabled, created lazily in src/net.cpp during model loading.
  • Custom allocators can be injected via ncnn::Option or per-extractor methods for specialized memory management such as DMA, NUMA-aware allocation, or memory tracing.

Frequently Asked Questions

Can I use the same allocator instance for both blob and workspace memory?

While technically possible, it is not recommended. The blob_allocator and workspace_allocator have fundamentally different lifetime patterns—persistent versus temporary. Using separate allocators allows the framework to optimize memory reuse and prevents fragmentation that would occur if temporary buffers competed with persistent tensors in the same pool. According to the ncnn source code in src/net.cpp, the framework explicitly maintains separate pool allocators for each type when use_local_pool_allocator is true.

How does use_local_pool_allocator affect performance?

When use_local_pool_allocator is true (the default), ncnn creates dedicated PoolAllocator instances for both allocator types in src/net.cpp (lines 1725-1732). These pool allocators cache memory blocks between forward passes, reducing system allocation overhead and improving cache locality. Disabling this forces every allocation to hit the system allocator via fastMalloc/fastFree, which can significantly degrade performance in latency-sensitive applications or when processing variable input sizes that trigger frequent reallocations.

Are Vulkan allocators compatible with CPU allocators?

No, VkBlobAllocator and VkWeightAllocator manage GPU device memory and are distinct from CPU allocators. They are set via blob_vkallocator and workspace_vkallocator in src/option.h (lines 48 and 51), whereas CPU allocators use blob_allocator and workspace_allocator. You cannot assign a CPU allocator to a GPU slot or vice versa without type errors, as they inherit from different base classes (VkAllocator vs Allocator) and handle fundamentally different memory spaces (device vs host).

When should I implement a custom allocator?

Implement a custom allocator when you need specialized memory management such as memory mapping, DMA buffer allocation, NUMA-aware allocation, or memory tracing. For example, embedded systems might require allocation from specific physical memory regions, or profiling tools might need to track allocation patterns to detect leaks. Inherit from ncnn::Allocator (defined in src/allocator.h) and override fastMalloc and fastFree, then inject your implementation via net.opt or per-extractor configuration before inference begins.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →