# Understanding one_blob_only and support_inplace Flags in ncnn for Memory-Efficient Inference

> Learn how ncnn's one_blob_only and support_inplace flags optimize memory for efficient inference by minimizing allocations and safely reusing buffers.

- Repository: [Tencent/ncnn](https://github.com/tencent/ncnn)
- Tags: internals
- Published: 2026-02-23

---

**The `one_blob_only` and `support_inplace` flags in ncnn's `Layer` class enable aggressive memory optimization by allowing the inference engine to bypass vector allocations for single-input layers and safely overwrite input buffers when the data is no longer needed.**

In the Tencent/ncnn deep learning framework—designed specifically for mobile and embedded devices—these two boolean flags declared in [`src/layer.h`](https://github.com/Tencent/ncnn/blob/main/src/layer.h) serve as critical hints to the forward execution engine in [`src/net.cpp`](https://github.com/Tencent/ncnn/blob/main/src/net.cpp). By signaling which layers accept exactly one input blob and which can compute results without allocating new memory, ncnn minimizes RAM usage and eliminates unnecessary data copies during neural network inference.

## What Are one_blob_only and support_inplace Flags?

Both flags are public members of the base `Layer` class defined at lines 45–50 of [`src/layer.h`](https://github.com/Tencent/ncnn/blob/main/src/layer.h):

- **`one_blob_only`** – Indicates the layer consumes exactly one input blob and produces exactly one output blob. When true, the framework can skip vector container overhead and reference blobs directly by index.
- **`support_inplace`** – Indicates the layer can safely compute its output by overwriting its input buffer. This is only valid when the input tensor's reference count is 1 (meaning no other layer needs that data).

Layer implementations set these flags in their constructors. For example, in [`src/layer/scale.cpp`](https://github.com/Tencent/ncnn/blob/main/src/layer/scale.cpp):

```cpp
Scale::Scale()
{
    one_blob_only = true;
    support_inplace = true;
}

```

## How These Flags Enable Single-Blob Fast Paths

When `one_blob_only` is true, the forward engine in [`src/net.cpp`](https://github.com/Tencent/ncnn/blob/main/src/net.cpp) takes an optimized code path that avoids std::vector allocations for `bottom_blobs` and `top_blobs`.

In `NetPrivate::do_forward_layer`, the engine checks this flag to determine how to package inputs:

```cpp
if (layer->one_blob_only)
{
    int bottom_blob_index = layer->bottoms[0];
    Mat& bottom_blob_ref = blob_mats[bottom_blob_index];
    // Direct reference to single blob, no vector construction
}

```

This optimization improves **CPU cache locality** by reducing pointer chasing through vector indirection and eliminates heap allocations during the inference hot path.

## In-Place Execution and Memory Optimization

The `support_inplace` flag drives ncnn's "light mode" (`opt.lightmode`) memory optimization strategy. When enabled, the engine attempts to reuse input buffers as output buffers, cutting peak memory usage significantly.

The logic in [`src/net.cpp`](https://github.com/Tencent/ncnn/blob/main/src/net.cpp) implements reference counting to ensure safety:

```cpp
if (opt.lightmode && layer->support_inplace)
{
    Mat bottom_blob;
    if (*bottom_blob_ref.refcount != 1)
    {
        // Blob is shared with other layers, must deep copy
        bottom_blob = bottom_blob_ref.clone(opt.blob_allocator);
    }
    else
    {
        // Safe to reuse buffer
        bottom_blob = bottom_blob_ref;
    }
    
    // Execute in-place computation
    int ret = layer->forward_inplace(bottom_blob, opt);
    blob_mats[layer->tops[0]] = bottom_blob;
}

```

**Key benefits of this approach:**
- **Zero-copy inference** when reference counts permit, eliminating `memcpy` operations
- **Reduced peak memory** footprint, critical for mobile GPUs and microcontrollers
- **Automatic safety fallback** to deep copies when blobs are shared across branches (e.g., residual connections)

## Implementation Details in the ncnn Source Code

These flags are defined in the base class at [`src/layer.h`](https://github.com/Tencent/ncnn/blob/main/src/layer.h) and consumed by the network executor at [`src/net.cpp`](https://github.com/Tencent/ncnn/blob/main/src/net.cpp). Specific layer implementations demonstrate the intended usage patterns:

- **[`src/layer/relu.cpp`](https://github.com/Tencent/ncnn/blob/main/src/layer/relu.cpp)** – Sets both flags to true because ReLU can transform data element-wise without allocating new memory
- **[`src/layer/reshape.cpp`](https://github.com/Tencent/ncnn/blob/main/src/layer/reshape.cpp)** – Sets `one_blob_only = true` but `support_inplace = false` because reshaping may change memory layout dimensions
- **[`src/layer/split.cpp`](https://github.com/Tencent/ncnn/blob/main/src/layer/split.cpp)** – Sets `one_blob_only = false` because it produces multiple output blobs from one input

The Vulkan GPU backend (`src/layer/vulkan`) respects these same flags when determining whether to use `forward_inplace` Vulkan pipelines or allocate new GPU buffers.

## Summary

- **`one_blob_only`** signals that a layer uses exactly one input and one output, enabling the ncnn engine to bypass vector allocations and take a fast path in [`src/net.cpp`](https://github.com/Tencent/ncnn/blob/main/src/net.cpp).
- **`support_inplace`** indicates that a layer can safely overwrite its input buffer, allowing the framework to reuse memory and avoid allocations when `opt.lightmode` is enabled.
- Together, these flags reduce **memory footprint**, eliminate unnecessary **data copies**, and improve **cache locality**—critical optimizations for deploying neural networks on mobile and embedded devices.

## Frequently Asked Questions

### What happens if a layer sets support_inplace to true but one_blob_only to false?

This combination is invalid and will cause runtime errors or undefined behavior. In-place computation requires exactly one input buffer to overwrite, so `support_inplace` should only be true when `one_blob_only` is also true. The ncnn engine assumes this invariant when calling `forward_inplace`.

### How does ncnn handle in-place layers when the input blob is shared across multiple branches?

The engine checks the reference count (`*bottom_blob_ref.refcount`) before executing in-place operations. If the count is greater than 1, indicating other layers still need the data, ncnn performs a deep copy via `clone()` to preserve the original buffer. This safety mechanism ensures correctness for residual connections and multi-branch networks.

### Can custom layers benefit from these optimization flags?

Yes. When implementing a custom layer by inheriting from `ncnn::Layer`, set `one_blob_only = true` in the constructor if your layer processes a single input tensor. Set `support_inplace = true` only if your `forward_inplace` implementation can safely overwrite the input buffer without affecting other layers. The network executor in [`src/net.cpp`](https://github.com/Tencent/ncnn/blob/main/src/net.cpp) will automatically apply the optimized paths when these flags are enabled.