ncnn Multi-Threaded Execution: How the openmp_blocktime Option Works

ncnn parallelizes CPU inference using OpenMP, and the openmp_blocktime option controls how long worker threads remain idle before sleeping to balance latency and power consumption.

Tencent's ncnn is a high-performance neural network inference framework optimized for mobile and desktop CPUs. By default, ncnn leverages OpenMP to distribute heavy computational kernels—such as convolutions and matrix multiplications—across multiple threads, while exposing fine-grained control over thread behavior through the openmp_blocktime configuration parameter.

How ncnn Implements Multi-Threaded Inference

ncnn achieves parallelism by annotating performance-critical loops with OpenMP directives. The library ships with a fallback implementation in src/simpleomp.cpp for builds without a real OpenMP runtime, though standard OpenMP libraries are used on most platforms.

Per-Layer Parallelization with OpenMP

Virtually every optimized layer kernel contains a pragma that distributes work across the thread pool defined in the global Option object. For example, x86-optimized layers use patterns like:

#pragma omp parallel for num_threads(opt.num_threads)

This pattern appears throughout the x86 layer implementations, including src/layer/x86/pooling_3x3_pack8.h and src/layer/x86/innerproduct_fp.h. The opt.num_threads value is drawn from the Option struct that users configure before constructing a Net.

Default Thread Pool Sizing

When creating an Option object, ncnn automatically sets the thread count to the number of physical "big" CPU cores detected by get_physical_big_cpu_count(). This default assignment occurs in src/option.cpp:

num_threads = get_physical_big_cpu_count();           // default

Users can override this default by assigning to opt.num_threads directly or by calling set_omp_num_threads() defined in src/cpu.cpp.

Understanding the openmp_blocktime Option

The openmp_blocktime member in the Option struct controls how long OpenMP worker threads remain in an idle spinning state before being parked (put to sleep). This setting directly impacts the trade-off between inference latency and power consumption.

The Role of Block Time in Thread Management

OpenMP runtimes (specifically LLVM/Clang and Intel implementations) maintain a pool of worker threads that persist after parallel regions complete. The block time determines the idle duration before a thread sleeps. A short block time reduces power consumption on bursty workloads, while a longer block time avoids the overhead of repeatedly waking threads for sequential inference calls.

ncnn exposes this tunable via Option::openmp_blocktime, which defaults to 20 milliseconds:

openmp_blocktime = 20;   // default = 20 ms

This default is defined in src/option.cpp.

How ncnn Applies Block Time Settings

To ensure the block time setting affects only ncnn inference without altering global application state, the Extractor class temporarily applies the user-defined value during inference and restores the original afterward. In src/net.cpp, the Extractor::extract() method implements this save-and-restore pattern:

int old_blocktime = get_kmp_blocktime();
set_kmp_blocktime(d->opt.openmp_blocktime);   // apply user option

...  // forward_layer calls that use OpenMP

set_kmp_blocktime(old_blocktime);             // restore

The actual implementation of set_kmp_blocktime() resides in src/cpu.cpp and forwards the call only when compiled with a KMP-compatible OpenMP runtime (Clang/LLVM or Intel):

void set_kmp_blocktime(int time_ms)
{
#if defined(_OPENMP) && (__clang__ || defined(_OPENMP_LLVM_RUNTIME))
    kmp_set_blocktime(time_ms);
#else
    (void)time_ms;   // no-op on other runtimes
#endif
}

On platforms using simpleomp.cpp (the fallback stub implementation), block time functions are no-ops, which is safe because these runtimes do not implement aggressive thread parking.

Configuring Multi-Threaded Execution in ncnn

You can configure both thread count and block time through the Option struct before loading your model.

Setting Thread Count and Block Time

The following example demonstrates configuring 8 threads with a 5ms block time:

#include "net.h"

int main()
{
    ncnn::Option opt;
    opt.num_threads      = 8;   // use 8 CPU threads
    opt.openmp_blocktime = 5;   // park idle threads after 5 ms

    ncnn::Net net;
    net.opt = opt;                     // apply the option to the network
    net.load_param("mobilenetv2.param");
    net.load_model("mobilenetv2.bin");

    ncnn::Mat in = ncnn::Mat::from_pixels_resize(...); // prepare input
    ncnn::Extractor ex = net.create_extractor();
    ex.input("data", in);

    ncnn::Mat out;
    ex.extract("prob", out);           // inference – block time applied automatically
}

All OpenMP handling is invisible to the user; the only knobs are opt.num_threads and opt.openmp_blocktime.

Querying Current Block Time

To inspect the current runtime block time value, include cpu.h and call:

int current = ncnn::get_kmp_blocktime();   // returns the value set in the runtime

This is useful for debugging or verifying that your configuration has taken effect.

Summary

  • ncnn multi-threaded execution relies on OpenMP pragmas (#pragma omp parallel for) inside layer kernels, using opt.num_threads to control parallelism.
  • The default thread count equals the number of physical big CPU cores as detected by get_physical_big_cpu_count() in src/option.cpp.
  • The openmp_blocktime option controls idle thread timeout, defaulting to 20ms to balance power efficiency and latency.
  • During Extractor::extract(), ncnn temporarily applies the configured block time via set_kmp_blocktime() and restores the original value afterward, ensuring changes remain local to the inference call.
  • Block time settings only affect KMP-compatible OpenMP runtimes (Clang/LLVM, Intel); other platforms use no-op stubs in src/simpleomp.cpp.

Frequently Asked Questions

What determines the default number of threads in ncnn?

By default, ncnn sets opt.num_threads to the count of physical "big" CPU cores using get_physical_big_cpu_count(), as implemented in src/option.cpp (lines 17–18). This prioritizes performance cores over efficiency cores on heterogeneous processors.

How does the openmp_blocktime setting affect inference performance?

A lower openmp_blocktime value (e.g., 0–5ms) reduces power consumption by parking threads quickly but may increase latency for back-to-back inferences due to thread wake-up overhead. A higher value (e.g., 20–200ms) keeps threads active, improving latency for bursty workloads at the cost of higher idle power consumption.

Can ncnn run multi-threaded without a real OpenMP library?

Yes. When compiled without OpenMP support, ncnn falls back to src/simpleomp.cpp, which provides single-threaded stubs for OpenMP functions. In this configuration, openmp_blocktime settings have no effect, and inference runs sequentially regardless of num_threads.

Is the openmp_blocktime setting global or local to each inference?

The setting is local to each inference call. According to the implementation in src/net.cpp (lines 38–42), Extractor::extract() saves the current runtime block time, applies the user-defined value for the duration of the forward pass, and restores the original value before returning. This prevents ncnn from permanently altering your application's OpenMP behavior.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →