# ncnn Multi-Threaded Execution: How the openmp_blocktime Option Works

> Discover how ncnn multi-threaded execution with OpenMP leverages openmp_blocktime to optimize worker thread idle time, balancing latency and power for efficient CPU inference.

- Repository: [Tencent/ncnn](https://github.com/tencent/ncnn)
- Tags: internals
- Published: 2026-02-23

---

**ncnn parallelizes CPU inference using OpenMP, and the `openmp_blocktime` option controls how long worker threads remain idle before sleeping to balance latency and power consumption.**

Tencent's ncnn is a high-performance neural network inference framework optimized for mobile and desktop CPUs. By default, ncnn leverages OpenMP to distribute heavy computational kernels—such as convolutions and matrix multiplications—across multiple threads, while exposing fine-grained control over thread behavior through the `openmp_blocktime` configuration parameter.

## How ncnn Implements Multi-Threaded Inference

ncnn achieves parallelism by annotating performance-critical loops with OpenMP directives. The library ships with a fallback implementation in [`src/simpleomp.cpp`](https://github.com/Tencent/ncnn/blob/main/src/simpleomp.cpp) for builds without a real OpenMP runtime, though standard OpenMP libraries are used on most platforms.

### Per-Layer Parallelization with OpenMP

Virtually every optimized layer kernel contains a pragma that distributes work across the thread pool defined in the global `Option` object. For example, x86-optimized layers use patterns like:

```cpp
#pragma omp parallel for num_threads(opt.num_threads)

```

This pattern appears throughout the x86 layer implementations, including [`src/layer/x86/pooling_3x3_pack8.h`](https://github.com/Tencent/ncnn/blob/main/src/layer/x86/pooling_3x3_pack8.h) and [`src/layer/x86/innerproduct_fp.h`](https://github.com/Tencent/ncnn/blob/main/src/layer/x86/innerproduct_fp.h). The `opt.num_threads` value is drawn from the `Option` struct that users configure before constructing a `Net`.

### Default Thread Pool Sizing

When creating an `Option` object, ncnn automatically sets the thread count to the number of physical "big" CPU cores detected by `get_physical_big_cpu_count()`. This default assignment occurs in [`src/option.cpp`](https://github.com/Tencent/ncnn/blob/main/src/option.cpp):

```cpp
num_threads = get_physical_big_cpu_count();           // default

```

Users can override this default by assigning to `opt.num_threads` directly or by calling `set_omp_num_threads()` defined in [`src/cpu.cpp`](https://github.com/Tencent/ncnn/blob/main/src/cpu.cpp).

## Understanding the openmp_blocktime Option

The `openmp_blocktime` member in the `Option` struct controls how long OpenMP worker threads remain in an idle spinning state before being parked (put to sleep). This setting directly impacts the trade-off between inference latency and power consumption.

### The Role of Block Time in Thread Management

OpenMP runtimes (specifically LLVM/Clang and Intel implementations) maintain a pool of worker threads that persist after parallel regions complete. The *block time* determines the idle duration before a thread sleeps. A short block time reduces power consumption on bursty workloads, while a longer block time avoids the overhead of repeatedly waking threads for sequential inference calls.

ncnn exposes this tunable via `Option::openmp_blocktime`, which defaults to **20 milliseconds**:

```cpp
openmp_blocktime = 20;   // default = 20 ms

```

This default is defined in [`src/option.cpp`](https://github.com/Tencent/ncnn/blob/main/src/option.cpp).

### How ncnn Applies Block Time Settings

To ensure the block time setting affects only ncnn inference without altering global application state, the `Extractor` class temporarily applies the user-defined value during inference and restores the original afterward. In [`src/net.cpp`](https://github.com/Tencent/ncnn/blob/main/src/net.cpp), the `Extractor::extract()` method implements this save-and-restore pattern:

```cpp
int old_blocktime = get_kmp_blocktime();
set_kmp_blocktime(d->opt.openmp_blocktime);   // apply user option

...  // forward_layer calls that use OpenMP

set_kmp_blocktime(old_blocktime);             // restore

```

The actual implementation of `set_kmp_blocktime()` resides in [`src/cpu.cpp`](https://github.com/Tencent/ncnn/blob/main/src/cpu.cpp) and forwards the call only when compiled with a KMP-compatible OpenMP runtime (Clang/LLVM or Intel):

```cpp
void set_kmp_blocktime(int time_ms)
{
#if defined(_OPENMP) && (__clang__ || defined(_OPENMP_LLVM_RUNTIME))
    kmp_set_blocktime(time_ms);
#else
    (void)time_ms;   // no-op on other runtimes
#endif
}

```

On platforms using [`simpleomp.cpp`](https://github.com/Tencent/ncnn/blob/main/simpleomp.cpp) (the fallback stub implementation), block time functions are no-ops, which is safe because these runtimes do not implement aggressive thread parking.

## Configuring Multi-Threaded Execution in ncnn

You can configure both thread count and block time through the `Option` struct before loading your model.

### Setting Thread Count and Block Time

The following example demonstrates configuring 8 threads with a 5ms block time:

```cpp
#include "net.h"

int main()
{
    ncnn::Option opt;
    opt.num_threads      = 8;   // use 8 CPU threads
    opt.openmp_blocktime = 5;   // park idle threads after 5 ms

    ncnn::Net net;
    net.opt = opt;                     // apply the option to the network
    net.load_param("mobilenetv2.param");
    net.load_model("mobilenetv2.bin");

    ncnn::Mat in = ncnn::Mat::from_pixels_resize(...); // prepare input
    ncnn::Extractor ex = net.create_extractor();
    ex.input("data", in);

    ncnn::Mat out;
    ex.extract("prob", out);           // inference – block time applied automatically
}

```

All OpenMP handling is invisible to the user; the only knobs are `opt.num_threads` and `opt.openmp_blocktime`.

### Querying Current Block Time

To inspect the current runtime block time value, include [`cpu.h`](https://github.com/Tencent/ncnn/blob/main/cpu.h) and call:

```cpp
int current = ncnn::get_kmp_blocktime();   // returns the value set in the runtime

```

This is useful for debugging or verifying that your configuration has taken effect.

## Summary

- **ncnn multi-threaded execution** relies on OpenMP pragmas (`#pragma omp parallel for`) inside layer kernels, using `opt.num_threads` to control parallelism.
- The default thread count equals the number of physical big CPU cores as detected by `get_physical_big_cpu_count()` in [`src/option.cpp`](https://github.com/Tencent/ncnn/blob/main/src/option.cpp).
- The `openmp_blocktime` option controls idle thread timeout, defaulting to **20ms** to balance power efficiency and latency.
- During `Extractor::extract()`, ncnn temporarily applies the configured block time via `set_kmp_blocktime()` and restores the original value afterward, ensuring changes remain local to the inference call.
- Block time settings only affect KMP-compatible OpenMP runtimes (Clang/LLVM, Intel); other platforms use no-op stubs in [`src/simpleomp.cpp`](https://github.com/Tencent/ncnn/blob/main/src/simpleomp.cpp).

## Frequently Asked Questions

### What determines the default number of threads in ncnn?

By default, ncnn sets `opt.num_threads` to the count of physical "big" CPU cores using `get_physical_big_cpu_count()`, as implemented in [`src/option.cpp`](https://github.com/Tencent/ncnn/blob/main/src/option.cpp) (lines 17–18). This prioritizes performance cores over efficiency cores on heterogeneous processors.

### How does the openmp_blocktime setting affect inference performance?

A lower `openmp_blocktime` value (e.g., 0–5ms) reduces power consumption by parking threads quickly but may increase latency for back-to-back inferences due to thread wake-up overhead. A higher value (e.g., 20–200ms) keeps threads active, improving latency for bursty workloads at the cost of higher idle power consumption.

### Can ncnn run multi-threaded without a real OpenMP library?

Yes. When compiled without OpenMP support, ncnn falls back to [`src/simpleomp.cpp`](https://github.com/Tencent/ncnn/blob/main/src/simpleomp.cpp), which provides single-threaded stubs for OpenMP functions. In this configuration, `openmp_blocktime` settings have no effect, and inference runs sequentially regardless of `num_threads`.

### Is the openmp_blocktime setting global or local to each inference?

The setting is local to each inference call. According to the implementation in [`src/net.cpp`](https://github.com/Tencent/ncnn/blob/main/src/net.cpp) (lines 38–42), `Extractor::extract()` saves the current runtime block time, applies the user-defined value for the duration of the forward pass, and restores the original value before returning. This prevents ncnn from permanently altering your application's OpenMP behavior.