# How ncnn's Extractor Class Handles Inference: Understanding `set_light_mode()` for Memory Reduction

> Discover how ncnn's Extractor class recycles buffers with set_light_mode() to slash memory usage during inference. Optimize your neural networks now!

- Repository: [Tencent/ncnn](https://github.com/tencent/ncnn)
- Tags: internals
- Published: 2026-02-23

---

**Enabling `set_light_mode()` allows ncnn's Extractor to recycle intermediate buffers via in-place operations and a local memory pool, significantly reducing peak memory usage during neural network inference.**

The Tencent/ncnn library provides a lightweight neural network inference framework optimized for mobile and embedded platforms. Central to its design is the `Extractor` class, which orchestrates the forward pass through a loaded network while managing tensor memory efficiently. Understanding how the Extractor handles computation—and how `set_light_mode()` optimizes memory allocation—helps developers maximize performance on resource-constrained devices.

## The ncnn Extractor Inference Pipeline

The `Extractor` class serves as the execution engine for neural network graphs loaded into an `ncnn::Net` instance. The inference process follows a structured lifecycle from initialization to output extraction.

### Creating the Extractor Instance

When you call `net.create_extractor()`, the framework instantiates an `Extractor` object through the constructor defined in `[src/net.cpp:2272-2277]`. This constructor performs two critical setup operations:

1. It allocates a `blob_mats` vector sized to the total number of blobs in the network graph
2. It copies the network's default `Option` configuration, including the `lightmode` flag

```cpp
ncnn::Net net;
net.load_param("model.param");
net.load_model("model.bin");
ncnn::Extractor ex = net.create_extractor();  // Creates extractor with default options

```

The Extractor maintains its own copy of inference options, allowing per-extraction configuration without affecting the parent network.

### Input Processing and Lazy Evaluation

Input tensors are registered via `Extractor::input(int blob_index, const Mat& in)` or its name-based overload. This method stores the provided matrix in `d->blob_mats[blob_index]` without triggering computation. The design follows a lazy evaluation pattern where no layer execution occurs until explicitly requested.

### Layer Execution and Output Extraction

The actual inference happens inside `Extractor::extract(int blob_index, Mat& feat, int type)`, implemented in `[src/net.cpp:2433-2509]`. When you request a specific output blob, the Extractor:

1. **Validates blob state** – Checks if the target blob has been materialized (dims == 0)
2. **Identifies the producer layer** – Looks up `layer_index = d->net->blobs()[blob_index].producer`
3. **Allocates temporary resources** – Uses `opt.use_local_pool_allocator` when enabled for intermediate tensors
4. **Executes forward pass** – Invokes the network's internal `forward_layer` routine, writing output to `d->blob_mats[blob_index]`
5. **Handles memory detachment** – When using the local pool allocator, clones the tensor (`feat = feat.clone()`) if it still points to the temporary allocator, allowing immediate pool reclamation

If Vulkan is enabled, this pipeline includes GPU-side execution and device-to-host transfer steps before returning the final `Mat` object.

## How `set_light_mode()` Reduces Memory Usage

The `set_light_mode(bool enable)` method, defined in `[src/net.cpp:2353-2356]`, controls the `opt.lightmode` flag within the Extractor's options. By default, `lightmode` is set to `true` as defined in `[src/option.cpp:10-13]`, enabling aggressive memory optimization strategies throughout inference.

### In-Place Buffer Reuse

When light mode is active, layer implementations check `if (opt.lightmode && layer->support_inplace)` before executing. This allows layers to write output directly over input buffers when the operation permits, eliminating redundant memory allocations for intermediate results. The Extractor coordinates these in-place operations to ensure data integrity while maximizing buffer utilization.

### Local Pool Allocator Strategy

Light mode activates the `local_blob_allocator`, a temporary memory pool dedicated to intermediate tensors during extraction. Unlike standard allocation patterns that persist buffers for the Extractor's lifetime, this approach:

- Allocates intermediate results from a reusable pool
- Returns memory to the pool immediately after blob consumption
- Reduces malloc/free fragmentation on both CPU and GPU (Vulkan) paths

The critical memory management logic appears in the extraction flow where the code verifies `if (d->opt.use_local_pool_allocator && feat.allocator == d->net->d->local_blob_allocator)` and performs cloning to detach from the temporary pool before returning results to the user.

### Impact on Peak Memory Footprint

Enabling light mode provides several concrete memory benefits:

- **Early network destruction** – Because intermediate results are owned by the Extractor's local pool rather than the network, the parent `Net` object can be destroyed while the Extractor remains active
- **Reduced peak RAM** – Temporary buffers are reclaimed as soon as downstream layers consume them, rather than accumulating throughout the forward pass
- **Deterministic reuse** – The pool allocator recycles the same memory blocks across multiple inference calls, avoiding system allocator overhead

When disabled (`set_light_mode(false)`), each layer allocates fresh output buffers that persist for the Extractor's lifetime, increasing memory consumption but preserving all intermediate tensors for debugging.

## Implementation Details in Source Code

### Option Structure and Defaults

The light mode behavior originates in the `Option` struct defined in [`src/option.h`](https://github.com/Tencent/ncnn/blob/main/src/option.h) with defaults set in `[src/option.cpp:10-13]`:

```cpp
Option::Option()
{
    lightmode = true;
    // ... other defaults
}

```

This default ensures that new Extractor instances automatically optimize for memory efficiency unless explicitly configured otherwise.

### Extractor Configuration Methods

The `set_light_mode()` implementation provides a simple interface to override the default behavior:

```cpp
void Extractor::set_light_mode(bool enable) { 
    d->opt.lightmode = enable; 
}

```

This modification affects all subsequent layer executions within that specific Extractor instance, allowing fine-grained control over memory strategies for different inference scenarios.

## Practical Usage Examples

### C++ Implementation

```cpp
#include "net.h"
#include "mat.h"

int main()
{
    ncnn::Net net;
    net.load_param("mobilenetv2.param");
    net.load_model("mobilenetv2.bin");

    // Create extractor with light mode enabled (default)
    ncnn::Extractor ex = net.create_extractor();
    ex.set_light_mode(true);  // Explicitly enable memory optimization
    
    // Prepare input
    ncnn::Mat in = ncnn::Mat::from_pixels_resize(
        image_data, ncnn::Mat::PIXEL_RGB, 
        img_w, img_h, input_w, input_h
    );
    
    ex.input(0, in);  // Set input blob by index
    
    ncnn::Mat out;
    ex.extract(128, out);  // Extract output from layer 128
    
    // out contains results with minimal memory footprint
    return 0;
}

```

### Python Bindings

```python
import ncnn

net = ncnn.Net()
net.load_param('mobilenetv2.param')
net.load_model('mobilenetv2.bin')

ex = net.create_extractor()
ex.set_light_mode(True)  # Enable light mode for memory efficiency

ex.input('data', input_mat)
ret, output_mat = ex.extract('prob')

```

Both examples demonstrate the standard workflow: load the network, create an Extractor, explicitly configure `set_light_mode()` for clarity, feed inputs, and extract results with optimized memory allocation.

## Summary

- **Lazy evaluation** – The Extractor defers computation until `extract()` is called, building a dynamic execution graph based on requested outputs
- **Light mode mechanics** – `set_light_mode(true)` enables in-place layer operations and temporary pool allocation via `[src/net.cpp:2353-2356]`
- **Memory optimization** – The local blob allocator reclaims intermediate buffers immediately after use, reducing peak memory usage compared to persistent allocation strategies
- **Default behavior** – Light mode is enabled by default according to `[src/option.cpp:10-13]`, but can be disabled for debugging scenarios requiring intermediate tensor inspection
- **Resource isolation** – Extractors maintain independent option states and can outlive their parent Net objects when using light mode memory strategies

## Frequently Asked Questions

### What happens if I disable light mode in ncnn's Extractor?

Disabling light mode via `set_light_mode(false)` forces the Extractor to allocate fresh output buffers for every layer and persist them for the Extractor's lifetime. While this increases memory consumption, it keeps all intermediate blobs accessible for debugging or layer-wise analysis, as none are recycled through the local pool allocator.

### Can I use set_light_mode() with Vulkan GPU inference?

Yes, light mode works with both CPU and Vulkan backends. When using Vulkan, the memory pool optimization applies to GPU allocations as well, reducing peak VRAM usage. The Extractor's cleanup routine in `[src/net.cpp]` releases locally-acquired Vulkan allocators when `clear()` is called or the Extractor is destroyed.

### Why does the Extractor clone tensors when using the local pool allocator?

According to the implementation in `[src/net.cpp:2433-2509]`, when `opt.use_local_pool_allocator` is active and an output tensor still points to the `local_blob_allocator`, the Extractor calls `feat.clone()` to copy data into standard memory. This detaches the result from the temporary pool, allowing the Extractor to reclaim those temporary buffers immediately while returning a safe, independent copy to the caller.

### Is light mode suitable for all neural network architectures?

Light mode benefits most standard feed-forward networks, but architectures requiring persistent intermediate states or residual connections that reuse earlier layer outputs may need careful validation. The `support_inplace` checks within individual layer implementations ensure safety, but complex models with custom layers should be tested with both modes to verify numerical correctness.