ncnn void* Parameter in ncnn::Mat: External Memory vs BlobMemoryPool
The void* parameter in ncnn::Mat enables zero-copy wrapping of external memory without reference counting, while ncnn::BlobMemoryPool provides managed, reference-counted allocation with automatic lifecycle management and block reuse.
The ncnn inference framework from Tencent uses ncnn::Mat as its core tensor container. Understanding the implications of the void* parameter in ncnn::Mat constructors—and how this differs from the ncnn::BlobMemoryPool allocation strategy—is critical for optimizing memory usage in high-performance inference pipelines.
What the void* Parameter Does in ncnn::Mat
ncnn::Mat provides constructors in src/mat.h (lines 58‑90) that accept an external memory pointer (void* data) instead of allocating their own buffer:
Mat(int w, void* data, size_t elemsize = 4u, Allocator* allocator = 0);
Mat(int w, int h, void* data, size_t elemsize = 4u, Allocator* allocator = 0);
Mat(int w, int h, int c, void* data, size_t elemsize = 4u, Allocator* allocator = 0);
Mat(int w, int h, int d, int c, void* data, size_t elemsize = 4u, Allocator* allocator = 0);
When using these overloads, ncnn::Mat does not own the memory. The caller must keep the buffer alive for the entire lifetime of the Mat object.
Implications of External Memory Wrapping
| Aspect | void* External Buffer |
Default Allocator Allocation |
|---|---|---|
| Ownership | Mat does not own memory; caller manages lifetime. |
Mat owns memory via reference counting. |
| Reference Counting | refcount is set to 0; addref()/release() do not manage the buffer. |
refcount is non‑null; enables safe shallow copies. |
| Allocator Usage | The allocator argument is ignored. |
Uses fastMalloc/fastFree via the supplied or default allocator. |
| Thread Safety | Caller must guarantee no concurrent modification or freeing. | ncnn allocators (e.g., BlobAllocator, VkBlobAllocator) are thread‑safe. |
| Alignment | Caller must ensure pointer satisfies alignSize(..., 16) requirements. |
Allocator guarantees 16‑byte alignment automatically. |
| Performance | Zero allocation overhead; ideal for zero‑copy with OpenCV or custom pools. | Slight overhead for allocation and reference counting, but benefits from memory‑pool reuse. |
Because the void* constructors intentionally set refcount to 0, ncnn treats the memory as user‑provided, enabling zero‑copy paths for interoperability with external libraries.
How ncnn::BlobMemoryPool Manages Memory
A Blob in ncnn represents the output of a layer and holds a Mat describing its shape. The actual tensor data is allocated from blob memory pools managed by the framework’s allocators.
Blob Structure and Network Integration
In src/blob.h (lines 20‑30), the Blob class is defined as:
class Blob {
public:
// ...
Mat shape; // shape only, no data initially
};
When a network is loaded and executed in src/net.cpp (lines 56, 1154‑1174):
- Net creates a vector of
Blobobjects (std::vector<Blob> blobs). - For each Blob requiring storage, the Net uses its BlobAllocator (stored inside the
Mat) to allocate memory viafastMalloc. - The allocator maintains memory blocks and buffer pools for reuse.
Allocator Architecture and Pooling
The base Allocator class in src/allocator.h defines the interface:
class Allocator {
public:
virtual void* fastMalloc(size_t size) = 0;
virtual void fastFree(void* ptr) = 0;
};
In src/allocator.cpp (lines 650‑695), the BlobAllocator implements pooling logic:
- Memory blocks are kept in a free-list (
buffer_blocks,buffer_pool). - When a new tensor is allocated, the allocator first checks for a reusable block of sufficient size.
- Freed memory returns to the pool rather than the OS, drastically reducing fragmentation and allocation overhead during inference.
For GPU inference, VkBlobAllocator in src/gpu.cpp (lines 3221‑3226) manages VkBufferMemory and VkImageMemory objects using the same pooling principles, enabling efficient reuse of GPU buffers across frames.
Key Differences: External void* vs BlobMemoryPool
| Feature | Mat with void* |
Blob Memory Pool (Allocator-backed Mat) |
|---|---|---|
| Memory Source | External user-provided buffer. | Internal pool via fastMalloc/fastFree. |
| Reference Counting | Disabled (refcount == 0). |
Enabled (refcount > 0). |
| Copy Semantics | Shallow copy shares raw pointer without ownership safety. | Shallow copy increments refcount, enabling safe sharing. |
| Deallocation | Caller must free manually. | Allocator frees automatically when refcount reaches zero. |
| Use Case | Interoperability with existing buffers, zero-copy pipelines. | Regular network inference where ncnn controls memory. |
| Performance Trade-off | No allocation cost, but requires careful lifetime management. | Slight allocation cost, but benefits from pooling and alignment. |
Practical Code Examples
Zero-Copy Input Using void*
When integrating with OpenCV or other external libraries, use the void* constructor to avoid data duplication:
// External image buffer from OpenCV (already allocated)
unsigned char* rgb_data = img.data; // img is cv::Mat with continuous memory
int w = img.cols;
int h = img.rows;
// Wrap external data without copying
ncnn::Mat in = ncnn::Mat(w, h, 3, rgb_data, 4u); // 4 bytes per pixel
net.input("data", in); // ncnn will NOT free rgb_data
// ... run inference ...
Critical: The caller must ensure rgb_data remains valid until inference completes and in is destroyed.
Regular Allocation via Blob Memory Pool
For standard inference where ncnn manages memory:
// Let ncnn allocate from its internal pool
ncnn::Mat in = ncnn::Mat(w, h, 3); // Uses default Allocator
// Fill with data (e.g., copy from external source)
memcpy(in.data, rgb_data, w * h * 3 * sizeof(float));
// Memory is reference-counted and automatically released
net.input("data", in);
net.extract("prob", out);
Here in’s data lives in the internal memory pool, is reference-counted, and will be automatically reclaimed when no longer referenced.
Using a Custom Allocator (Advanced)
You can provide a custom allocator while still benefiting from ncnn's pooling infrastructure:
class MyAllocator : public ncnn::Allocator {
public:
virtual void* fastMalloc(size_t size) override {
// Custom allocation from a pre-allocated arena
return my_arena.alloc(size);
}
virtual void fastFree(void* ptr) override {
// Return to arena or defer cleanup
}
};
MyAllocator my_alloc;
ncnn::Mat a = ncnn::Mat(64, 64, 3, &my_alloc); // Uses custom allocator
When an Allocator* is supplied, the Blob memory pool mechanism (reference counting, alignment) still applies, but physical memory comes from the user-provided allocator.
Summary
- The
void*constructors inncnn::Mat(defined insrc/mat.h) enable zero-copy integration with external buffers by disabling reference counting (refcount == 0), placing full lifetime responsibility on the caller. - BlobMemoryPool (implemented via
Allocatorsubclasses insrc/allocator.cpp) provides reference-counted, aligned memory with automatic pooling and reuse, suitable for standard inference workflows. - Use
void*only when you require zero-copy interoperability with existing memory (e.g., OpenCVcv::Mat) and can guarantee buffer stability throughout inference. - Use default allocation (or custom
Allocatorsubclasses) for regular tensor operations to benefit from ncnn's sophisticated memory pooling, alignment guarantees, and automatic lifecycle management.
Frequently Asked Questions
What happens if I free the external buffer while ncnn::Mat still references it?
If you allocate memory externally and pass it to ncnn::Mat via the void* constructor, then free that memory before the Mat is destroyed or used, you will encounter undefined behavior (typically a segmentation fault or data corruption). Because the refcount is set to 0, ncnn assumes it does not own the memory and will not attempt to prevent access or delay cleanup. You must ensure the external buffer outlives all ncnn::Mat instances that reference it.
Can I mix void* Mats with BlobMemoryPool allocations in the same network?
Yes, you can use both approaches simultaneously. For example, you might wrap input image data from an external source using the void* constructor to avoid copying, while allowing intermediate layer outputs and final results to be allocated from the BlobMemoryPool. The network execution logic in src/net.cpp handles each Mat according to its internal state—those with refcount == 0 (external) are left untouched, while those with active reference counts are managed through the allocator's pooling mechanism.
Does using the void* constructor affect GPU inference with Vulkan?
The void* constructor works with CPU memory (Mat data). For Vulkan GPU inference, memory is managed by VkBlobAllocator and VkImageAllocator classes (defined in src/gpu.cpp and src/allocator.h). While you can upload CPU data wrapped via void* to GPU using ncnn::VkCompute transfer commands, the GPU-side storage itself is always allocated through the Vulkan blob memory pool. You cannot directly pass a void* pointer to GPU memory—GPU buffers require driver-managed allocation through the VkBlobAllocator infrastructure.
How do I ensure proper memory alignment when using the void* parameter?
When using the void* constructors, you must guarantee that the external pointer satisfies ncnn's alignment requirements. According to src/allocator.h, ncnn typically requires 16-byte alignment (using alignSize(size, 16)). If your external buffer comes from malloc or new, it may not meet these requirements, potentially causing crashes on SIMD-optimized operations. Use platform-specific aligned allocation (e.g., _aligned_malloc on Windows, posix_memalign on Linux) or ensure your buffer is allocated via ncnn's own fastMalloc before wrapping it with the void* constructor.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →