# ncnn Memory Allocation Strategies: Understanding blob_allocator vs workspace_allocator in the Option Class

> Discover ncnn memory allocation strategies! Learn the key differences between blob_allocator for persistent layer outputs and workspace_allocator for temporary computation buffers in the Option class.

- Repository: [Tencent/ncnn](https://github.com/tencent/ncnn)
- Tags: internals
- Published: 2026-02-23

---

**In ncnn, `blob_allocator` manages persistent memory for layer outputs that survive across forward passes, while `workspace_allocator` provides temporary scratch buffers used only during single-layer computation and immediately recycled.**

The Tencent/ncnn inference framework separates memory management into two distinct strategies through the **`Option`** class. Understanding the difference between `blob_allocator` and `workspace_allocator` is essential for optimizing memory usage in deep learning deployments, particularly when implementing custom allocators or deploying on memory-constrained edge devices.

## Core Differences Between blob_allocator and workspace_allocator

### Purpose and Memory Lifetime

The fundamental distinction lies in data persistence:

- **`blob_allocator`**: Allocates **persistent** storage for layer outputs (blobs). These tensors survive across layer boundaries and multiple forward passes until explicitly released when the network is cleared or the next forward pass begins.
- **`workspace_allocator`**: Provides **temporary** scratch space for intermediate calculations within a single layer. Examples include `im2col` transformations, Winograd tiles, or temporary GEMM buffers. This memory is allocated, used, and freed within the same forward call.

### Default PoolAllocator Behavior

When `opt.use_local_pool_allocator` is `true` (the default) and no custom allocator is set, ncnn lazily creates separate **`PoolAllocator`** instances in `Net::load_model` within [`src/net.cpp`](https://github.com/Tencent/ncnn/blob/main/src/net.cpp) (lines 1725-1732). The framework maintains `d->local_blob_allocator` for persistent data and `d->local_workspace_allocator` for temporary buffers, ensuring memory reuse without fragmentation between inference passes.

## Implementation in Layer Code

### Allocating Output Blobs

Layers allocate persistent output tensors using `blob_allocator` via `Mat::create`. In [`src/layer/x86/slice_x86.cpp`](https://github.com/Tencent/ncnn/blob/main/src/layer/x86/slice_x86.cpp) at line 78:

```cpp
top_blob.create(out_w, out_h, out_c, opt.blob_allocator);

```

This allocates memory that persists beyond the `slice_x86` layer's execution, holding the sliced output for subsequent layers.

### Allocating Workspace Buffers

For temporary computation buffers, layers use `workspace_allocator`. In [`src/layer/x86/convolution_im2col_gemm.h`](https://github.com/Tencent/ncnn/blob/main/src/layer/x86/convolution_im2col_gemm.h) at line 4615:

```cpp
Mat BT(..., opt.workspace_allocator);

```

This buffer exists only during the convolution operation. When the `Mat` goes out of scope at the end of the forward function, the memory returns to the workspace pool for immediate reuse by the next layer.

## Configuring Custom Allocators

### Implementing a Custom Allocator

Create a class inheriting from `ncnn::Allocator` and override `fastMalloc` and `fastFree` as defined in [`src/allocator.h`](https://github.com/Tencent/ncnn/blob/main/src/allocator.h):

```cpp
class MyAllocator : public ncnn::Allocator {
public:
    void* fastMalloc(size_t size) override {
        // Custom allocation logic (e.g., DMA, pinned memory)
        return malloc(size);
    }
    void fastFree(void* ptr) override {
        // Custom deallocation logic
        free(ptr);
    }
};

```

### Setting Allocators on Net and Extractor

Configure allocators before loading the model to control memory placement:

```cpp
ncnn::Net net;
MyAllocator blob_alloc;
MyAllocator workspace_alloc;

net.opt.blob_allocator = &blob_alloc;
net.opt.workspace_allocator = &workspace_alloc;
net.opt.use_local_pool_allocator = false;  // Disable default pool

net.load_model("model.bin");

```

Alternatively, set per-extractor at runtime for specific inference contexts:

```cpp
auto ex = net.create_extractor();
ex.set_blob_allocator(&blob_alloc);
ex.set_workspace_allocator(&workspace_alloc);

```

### Vulkan GPU Allocators

For GPU inference, use the Vulkan-specific variants declared in [`src/option.h`](https://github.com/Tencent/ncnn/blob/main/src/option.h) (lines 48 and 51):

```cpp
ncnn::VulkanDevice* vkdev = ncnn::get_gpu_device(0);
ncnn::VkBlobAllocator vk_blob_alloc(vkdev);
ncnn::VkWeightAllocator vk_workspace_alloc(vkdev);  // Can serve as workspace

net.opt.blob_vkallocator = &vk_blob_alloc;
net.opt.workspace_vkallocator = &vk_workspace_alloc;

```

## Summary

- **`blob_allocator`** manages **persistent** memory for layer outputs that survive across forward passes and network boundaries, allocated via `Mat::create` in layer implementations.
- **`workspace_allocator`** provides **temporary** scratch buffers used only during single-layer computation (e.g., `im2col`, Winograd transforms) and immediately recycled.
- Both default to separate **`PoolAllocator`** instances when `use_local_pool_allocator` is enabled, created lazily in [`src/net.cpp`](https://github.com/Tencent/ncnn/blob/main/src/net.cpp) during model loading.
- Custom allocators can be injected via `ncnn::Option` or per-extractor methods for specialized memory management such as DMA, NUMA-aware allocation, or memory tracing.

## Frequently Asked Questions

### Can I use the same allocator instance for both blob and workspace memory?

While technically possible, it is **not recommended**. The `blob_allocator` and `workspace_allocator` have fundamentally different lifetime patterns—persistent versus temporary. Using separate allocators allows the framework to optimize memory reuse and prevents fragmentation that would occur if temporary buffers competed with persistent tensors in the same pool. According to the ncnn source code in [`src/net.cpp`](https://github.com/Tencent/ncnn/blob/main/src/net.cpp), the framework explicitly maintains separate pool allocators for each type when `use_local_pool_allocator` is true.

### How does use_local_pool_allocator affect performance?

When `use_local_pool_allocator` is true (the default), ncnn creates dedicated **`PoolAllocator`** instances for both allocator types in [`src/net.cpp`](https://github.com/Tencent/ncnn/blob/main/src/net.cpp) (lines 1725-1732). These pool allocators cache memory blocks between forward passes, reducing system allocation overhead and improving cache locality. Disabling this forces every allocation to hit the system allocator via `fastMalloc`/`fastFree`, which can significantly degrade performance in latency-sensitive applications or when processing variable input sizes that trigger frequent reallocations.

### Are Vulkan allocators compatible with CPU allocators?

No, **`VkBlobAllocator`** and **`VkWeightAllocator`** manage GPU device memory and are distinct from CPU allocators. They are set via `blob_vkallocator` and `workspace_vkallocator` in [`src/option.h`](https://github.com/Tencent/ncnn/blob/main/src/option.h) (lines 48 and 51), whereas CPU allocators use `blob_allocator` and `workspace_allocator`. You cannot assign a CPU allocator to a GPU slot or vice versa without type errors, as they inherit from different base classes (`VkAllocator` vs `Allocator`) and handle fundamentally different memory spaces (device vs host).

### When should I implement a custom allocator?

Implement a custom allocator when you need specialized memory management such as **memory mapping**, **DMA buffer allocation**, **NUMA-aware allocation**, or **memory tracing**. For example, embedded systems might require allocation from specific physical memory regions, or profiling tools might need to track allocation patterns to detect leaks. Inherit from `ncnn::Allocator` (defined in [`src/allocator.h`](https://github.com/Tencent/ncnn/blob/main/src/allocator.h)) and override `fastMalloc` and `fastFree`, then inject your implementation via `net.opt` or per-extractor configuration before inference begins.