Understanding one_blob_only and support_inplace Flags in ncnn for Memory-Efficient Inference
The one_blob_only and support_inplace flags in ncnn's Layer class enable aggressive memory optimization by allowing the inference engine to bypass vector allocations for single-input layers and safely overwrite input buffers when the data is no longer needed.
In the Tencent/ncnn deep learning framework—designed specifically for mobile and embedded devices—these two boolean flags declared in src/layer.h serve as critical hints to the forward execution engine in src/net.cpp. By signaling which layers accept exactly one input blob and which can compute results without allocating new memory, ncnn minimizes RAM usage and eliminates unnecessary data copies during neural network inference.
What Are one_blob_only and support_inplace Flags?
Both flags are public members of the base Layer class defined at lines 45–50 of src/layer.h:
one_blob_only– Indicates the layer consumes exactly one input blob and produces exactly one output blob. When true, the framework can skip vector container overhead and reference blobs directly by index.support_inplace– Indicates the layer can safely compute its output by overwriting its input buffer. This is only valid when the input tensor's reference count is 1 (meaning no other layer needs that data).
Layer implementations set these flags in their constructors. For example, in src/layer/scale.cpp:
Scale::Scale()
{
one_blob_only = true;
support_inplace = true;
}
How These Flags Enable Single-Blob Fast Paths
When one_blob_only is true, the forward engine in src/net.cpp takes an optimized code path that avoids std::vector allocations for bottom_blobs and top_blobs.
In NetPrivate::do_forward_layer, the engine checks this flag to determine how to package inputs:
if (layer->one_blob_only)
{
int bottom_blob_index = layer->bottoms[0];
Mat& bottom_blob_ref = blob_mats[bottom_blob_index];
// Direct reference to single blob, no vector construction
}
This optimization improves CPU cache locality by reducing pointer chasing through vector indirection and eliminates heap allocations during the inference hot path.
In-Place Execution and Memory Optimization
The support_inplace flag drives ncnn's "light mode" (opt.lightmode) memory optimization strategy. When enabled, the engine attempts to reuse input buffers as output buffers, cutting peak memory usage significantly.
The logic in src/net.cpp implements reference counting to ensure safety:
if (opt.lightmode && layer->support_inplace)
{
Mat bottom_blob;
if (*bottom_blob_ref.refcount != 1)
{
// Blob is shared with other layers, must deep copy
bottom_blob = bottom_blob_ref.clone(opt.blob_allocator);
}
else
{
// Safe to reuse buffer
bottom_blob = bottom_blob_ref;
}
// Execute in-place computation
int ret = layer->forward_inplace(bottom_blob, opt);
blob_mats[layer->tops[0]] = bottom_blob;
}
Key benefits of this approach:
- Zero-copy inference when reference counts permit, eliminating
memcpyoperations - Reduced peak memory footprint, critical for mobile GPUs and microcontrollers
- Automatic safety fallback to deep copies when blobs are shared across branches (e.g., residual connections)
Implementation Details in the ncnn Source Code
These flags are defined in the base class at src/layer.h and consumed by the network executor at src/net.cpp. Specific layer implementations demonstrate the intended usage patterns:
src/layer/relu.cpp– Sets both flags to true because ReLU can transform data element-wise without allocating new memorysrc/layer/reshape.cpp– Setsone_blob_only = truebutsupport_inplace = falsebecause reshaping may change memory layout dimensionssrc/layer/split.cpp– Setsone_blob_only = falsebecause it produces multiple output blobs from one input
The Vulkan GPU backend (src/layer/vulkan) respects these same flags when determining whether to use forward_inplace Vulkan pipelines or allocate new GPU buffers.
Summary
one_blob_onlysignals that a layer uses exactly one input and one output, enabling the ncnn engine to bypass vector allocations and take a fast path insrc/net.cpp.support_inplaceindicates that a layer can safely overwrite its input buffer, allowing the framework to reuse memory and avoid allocations whenopt.lightmodeis enabled.- Together, these flags reduce memory footprint, eliminate unnecessary data copies, and improve cache locality—critical optimizations for deploying neural networks on mobile and embedded devices.
Frequently Asked Questions
What happens if a layer sets support_inplace to true but one_blob_only to false?
This combination is invalid and will cause runtime errors or undefined behavior. In-place computation requires exactly one input buffer to overwrite, so support_inplace should only be true when one_blob_only is also true. The ncnn engine assumes this invariant when calling forward_inplace.
How does ncnn handle in-place layers when the input blob is shared across multiple branches?
The engine checks the reference count (*bottom_blob_ref.refcount) before executing in-place operations. If the count is greater than 1, indicating other layers still need the data, ncnn performs a deep copy via clone() to preserve the original buffer. This safety mechanism ensures correctness for residual connections and multi-branch networks.
Can custom layers benefit from these optimization flags?
Yes. When implementing a custom layer by inheriting from ncnn::Layer, set one_blob_only = true in the constructor if your layer processes a single input tensor. Set support_inplace = true only if your forward_inplace implementation can safely overwrite the input buffer without affecting other layers. The network executor in src/net.cpp will automatically apply the optimized paths when these flags are enabled.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →