How ncnn's Extractor Class Handles Inference: Understanding `set_light_mode()` for Memory Reduction

Enabling set_light_mode() allows ncnn's Extractor to recycle intermediate buffers via in-place operations and a local memory pool, significantly reducing peak memory usage during neural network inference.

The Tencent/ncnn library provides a lightweight neural network inference framework optimized for mobile and embedded platforms. Central to its design is the Extractor class, which orchestrates the forward pass through a loaded network while managing tensor memory efficiently. Understanding how the Extractor handles computation—and how set_light_mode() optimizes memory allocation—helps developers maximize performance on resource-constrained devices.

The ncnn Extractor Inference Pipeline

The Extractor class serves as the execution engine for neural network graphs loaded into an ncnn::Net instance. The inference process follows a structured lifecycle from initialization to output extraction.

Creating the Extractor Instance

When you call net.create_extractor(), the framework instantiates an Extractor object through the constructor defined in [src/net.cpp:2272-2277]. This constructor performs two critical setup operations:

  1. It allocates a blob_mats vector sized to the total number of blobs in the network graph
  2. It copies the network's default Option configuration, including the lightmode flag
ncnn::Net net;
net.load_param("model.param");
net.load_model("model.bin");
ncnn::Extractor ex = net.create_extractor();  // Creates extractor with default options

The Extractor maintains its own copy of inference options, allowing per-extraction configuration without affecting the parent network.

Input Processing and Lazy Evaluation

Input tensors are registered via Extractor::input(int blob_index, const Mat& in) or its name-based overload. This method stores the provided matrix in d->blob_mats[blob_index] without triggering computation. The design follows a lazy evaluation pattern where no layer execution occurs until explicitly requested.

Layer Execution and Output Extraction

The actual inference happens inside Extractor::extract(int blob_index, Mat& feat, int type), implemented in [src/net.cpp:2433-2509]. When you request a specific output blob, the Extractor:

  1. Validates blob state – Checks if the target blob has been materialized (dims == 0)
  2. Identifies the producer layer – Looks up layer_index = d->net->blobs()[blob_index].producer
  3. Allocates temporary resources – Uses opt.use_local_pool_allocator when enabled for intermediate tensors
  4. Executes forward pass – Invokes the network's internal forward_layer routine, writing output to d->blob_mats[blob_index]
  5. Handles memory detachment – When using the local pool allocator, clones the tensor (feat = feat.clone()) if it still points to the temporary allocator, allowing immediate pool reclamation

If Vulkan is enabled, this pipeline includes GPU-side execution and device-to-host transfer steps before returning the final Mat object.

How set_light_mode() Reduces Memory Usage

The set_light_mode(bool enable) method, defined in [src/net.cpp:2353-2356], controls the opt.lightmode flag within the Extractor's options. By default, lightmode is set to true as defined in [src/option.cpp:10-13], enabling aggressive memory optimization strategies throughout inference.

In-Place Buffer Reuse

When light mode is active, layer implementations check if (opt.lightmode && layer->support_inplace) before executing. This allows layers to write output directly over input buffers when the operation permits, eliminating redundant memory allocations for intermediate results. The Extractor coordinates these in-place operations to ensure data integrity while maximizing buffer utilization.

Local Pool Allocator Strategy

Light mode activates the local_blob_allocator, a temporary memory pool dedicated to intermediate tensors during extraction. Unlike standard allocation patterns that persist buffers for the Extractor's lifetime, this approach:

  • Allocates intermediate results from a reusable pool
  • Returns memory to the pool immediately after blob consumption
  • Reduces malloc/free fragmentation on both CPU and GPU (Vulkan) paths

The critical memory management logic appears in the extraction flow where the code verifies if (d->opt.use_local_pool_allocator && feat.allocator == d->net->d->local_blob_allocator) and performs cloning to detach from the temporary pool before returning results to the user.

Impact on Peak Memory Footprint

Enabling light mode provides several concrete memory benefits:

  • Early network destruction – Because intermediate results are owned by the Extractor's local pool rather than the network, the parent Net object can be destroyed while the Extractor remains active
  • Reduced peak RAM – Temporary buffers are reclaimed as soon as downstream layers consume them, rather than accumulating throughout the forward pass
  • Deterministic reuse – The pool allocator recycles the same memory blocks across multiple inference calls, avoiding system allocator overhead

When disabled (set_light_mode(false)), each layer allocates fresh output buffers that persist for the Extractor's lifetime, increasing memory consumption but preserving all intermediate tensors for debugging.

Implementation Details in Source Code

Option Structure and Defaults

The light mode behavior originates in the Option struct defined in src/option.h with defaults set in [src/option.cpp:10-13]:

Option::Option()
{
    lightmode = true;
    // ... other defaults
}

This default ensures that new Extractor instances automatically optimize for memory efficiency unless explicitly configured otherwise.

Extractor Configuration Methods

The set_light_mode() implementation provides a simple interface to override the default behavior:

void Extractor::set_light_mode(bool enable) { 
    d->opt.lightmode = enable; 
}

This modification affects all subsequent layer executions within that specific Extractor instance, allowing fine-grained control over memory strategies for different inference scenarios.

Practical Usage Examples

C++ Implementation

#include "net.h"
#include "mat.h"

int main()
{
    ncnn::Net net;
    net.load_param("mobilenetv2.param");
    net.load_model("mobilenetv2.bin");

    // Create extractor with light mode enabled (default)
    ncnn::Extractor ex = net.create_extractor();
    ex.set_light_mode(true);  // Explicitly enable memory optimization
    
    // Prepare input
    ncnn::Mat in = ncnn::Mat::from_pixels_resize(
        image_data, ncnn::Mat::PIXEL_RGB, 
        img_w, img_h, input_w, input_h
    );
    
    ex.input(0, in);  // Set input blob by index
    
    ncnn::Mat out;
    ex.extract(128, out);  // Extract output from layer 128
    
    // out contains results with minimal memory footprint
    return 0;
}

Python Bindings

import ncnn

net = ncnn.Net()
net.load_param('mobilenetv2.param')
net.load_model('mobilenetv2.bin')

ex = net.create_extractor()
ex.set_light_mode(True)  # Enable light mode for memory efficiency

ex.input('data', input_mat)
ret, output_mat = ex.extract('prob')

Both examples demonstrate the standard workflow: load the network, create an Extractor, explicitly configure set_light_mode() for clarity, feed inputs, and extract results with optimized memory allocation.

Summary

  • Lazy evaluation – The Extractor defers computation until extract() is called, building a dynamic execution graph based on requested outputs
  • Light mode mechanics – set_light_mode(true) enables in-place layer operations and temporary pool allocation via [src/net.cpp:2353-2356]
  • Memory optimization – The local blob allocator reclaims intermediate buffers immediately after use, reducing peak memory usage compared to persistent allocation strategies
  • Default behavior – Light mode is enabled by default according to [src/option.cpp:10-13], but can be disabled for debugging scenarios requiring intermediate tensor inspection
  • Resource isolation – Extractors maintain independent option states and can outlive their parent Net objects when using light mode memory strategies

Frequently Asked Questions

What happens if I disable light mode in ncnn's Extractor?

Disabling light mode via set_light_mode(false) forces the Extractor to allocate fresh output buffers for every layer and persist them for the Extractor's lifetime. While this increases memory consumption, it keeps all intermediate blobs accessible for debugging or layer-wise analysis, as none are recycled through the local pool allocator.

Can I use set_light_mode() with Vulkan GPU inference?

Yes, light mode works with both CPU and Vulkan backends. When using Vulkan, the memory pool optimization applies to GPU allocations as well, reducing peak VRAM usage. The Extractor's cleanup routine in [src/net.cpp] releases locally-acquired Vulkan allocators when clear() is called or the Extractor is destroyed.

Why does the Extractor clone tensors when using the local pool allocator?

According to the implementation in [src/net.cpp:2433-2509], when opt.use_local_pool_allocator is active and an output tensor still points to the local_blob_allocator, the Extractor calls feat.clone() to copy data into standard memory. This detaches the result from the temporary pool, allowing the Extractor to reclaim those temporary buffers immediately while returning a safe, independent copy to the caller.

Is light mode suitable for all neural network architectures?

Light mode benefits most standard feed-forward networks, but architectures requiring persistent intermediate states or residual connections that reuse earlier layer outputs may need careful validation. The support_inplace checks within individual layer implementations ensure safety, but complex models with custom layers should be tested with both modes to verify numerical correctness.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →