# Understanding the flush_denormals Option in ncnn for ARM Performance Optimization

> Optimize ARM performance with ncnn by understanding flush denormals. Learn how this option prevents costly micro-code paths on processors like Cortex A53 and A55.

- Repository: [Tencent/ncnn](https://github.com/tencent/ncnn)
- Tags: performance
- Published: 2026-02-23

---

**The `flush_denormals` option in ncnn controls how denormal floating-point values are handled, with mode 3 (default) enabling both Denormals-Are-Zero and Flush-To-Zero to prevent costly micro-code paths on ARM processors like Cortex-A53 and Cortex-A55.**

The Tencent/ncnn inference framework provides the `flush_denormals` configuration to optimize floating-point performance across different CPU architectures. This option is particularly critical when deploying neural networks on ARM-based mobile devices, where subnormal number handling can severely impact real-time inference speed.

## What Is the flush_denormals Option in ncnn?

In ncnn, `flush_denormals` is a field within the **Option** structure defined in [`src/option.h`](https://github.com/Tencent/ncnn/blob/main/src/option.h) (lines 115-121). It controls the processor's behavior regarding *denormal* (or subnormal) floating-point values—numbers smaller than the normal range that require special handling.

### The Four Operating Modes

The option accepts four integer values that map to specific hardware behaviors:

- **0** – DAZ off, FTZ off: Denormals are processed normally (slowest mode).
- **1** – DAZ on, FTZ off: Denormals are replaced with zero when read, but operation results are not forced to zero.
- **2** – DAZ off, FTZ on: Results of operations that would produce denormals are flushed to zero.
- **3** – DAZ on, FTZ on: Both modes enabled—the most aggressive flushing (default in ncnn).

## How ncnn Implements flush_denormals

The implementation spans multiple source files, handling state management differently across x86 and ARM architectures.

### State Management in net.cpp

Before each inference run, ncnn saves the existing floating-point state and applies the configured flush mode. In [`src/net.cpp`](https://github.com/Tencent/ncnn/blob/main/src/net.cpp) (lines 2441-2443), the code calls `get_flush_denormals` to preserve the previous state, then sets the new mode for the duration of the extraction. After inference completes, the original state is restored to avoid side effects on the calling application.

### x86 Implementation

On x86 platforms, ncnn utilizes SSE intrinsics to control hardware floating-point behavior. The implementation in [`src/cpu.cpp`](https://github.com/Tencent/ncnn/blob/main/src/cpu.cpp) (lines 3240-3267) uses `_MM_SET_DENORMALS_ZERO_MODE` and `_MM_SET_FLUSH_ZERO_MODE` to toggle DAZ and FTZ states when the option values 1-3 are selected.

### ARM NEON Implementation

ARM processors require explicit handling because hardware support for DAZ/FTZ varies across cores. Rather than relying solely on hardware flags, ncnn embeds flush-to-zero logic directly in NEON kernels. For example, in [`src/layer/arm/neon_mathfun.h`](https://github.com/Tencent/ncnn/blob/main/src/layer/arm/neon_mathfun.h) (line 63), the logarithm implementation forces denormals to zero:

```cpp
x = vmaxq_f32(x, vdupq_n_f32(0));  // force flush to zero on denormal values

```

This explicit clamping ensures consistent performance regardless of the ARM core's hardware floating-point capabilities.

## Why flush_denormals Matters for ARM Processor Performance

The default value of 3 (both DAZ and FTZ enabled) is not arbitrary—it addresses specific performance pathologies in ARM microarchitectures.

### The Cost of Denormal Values on ARM

Denormal floating-point values trigger **micro-code assist paths** on many ARM cores, particularly the widely deployed Cortex-A53 and Cortex-A55. When a calculation produces a subnormal result, the CPU must exit its standard fast-path pipeline and execute specialized microcode that can consume **dozens of additional cycles** per operation. In deep learning inference, where billions of floating-point operations execute per second, even occasional denormals can cause significant throughput drops.

### Impact on NEON SIMD Kernels

ncnn relies heavily on **NEON SIMD** vectorization for compute-intensive operators like convolution and matrix multiplication. SIMD instructions process multiple data elements simultaneously; if any lane produces a denormal, the entire vector operation may incur the slow path penalty. By forcing denormals to zero through the `flush_denormals` option and explicit NEON clamping, ncnn ensures that all arithmetic stays within the fast normal-range domain, maintaining consistent vector throughput.

### Default Optimization for A53/A55 Cores

Recognizing the prevalence of these cores in mobile devices, ncnn automatically enables aggressive denormal flushing for them. In [`src/option.cpp`](https://github.com/Tencent/ncnn/blob/main/src/option.cpp) (lines 68-70), the library sets `opt.flush_denormals = 3` by default when targeting Cortex-A53 or A55 architectures, ensuring optimal performance out-of-the-box without requiring manual configuration.

## Practical Usage Examples

Configure the option before loading your model to ensure the setting applies throughout the network lifecycle:

```cpp
ncnn::Option opt;
opt.flush_denormals = 3;          // Default: enable both DAZ and FTZ

ncnn::Net net;
net.opt = opt;                    // Apply before loading models
net.load_param("mobilenet_v2.param");
net.load_model("mobilenet_v2.bin");

```

To profile the performance impact of denormal handling or debug numerical issues, temporarily disable flushing:

```cpp
opt.flush_denormals = 0;          // Turn off DAZ/FTZ for debugging
net.opt = opt;

```

## Summary

- The `flush_denormals` option in ncnn controls hardware and software handling of subnormal floating-point values through four modes (0-3), with mode 3 (DAZ + FTZ) as the default.
- In [`src/net.cpp`](https://github.com/Tencent/ncnn/blob/main/src/net.cpp), ncnn saves and restores the floating-point state around inference runs to apply the setting without side effects.
- x86 platforms use SSE intrinsics in [`src/cpu.cpp`](https://github.com/Tencent/ncnn/blob/main/src/cpu.cpp) to toggle hardware DAZ/FTZ modes, while ARM platforms rely on explicit NEON clamping in kernels like [`src/layer/arm/neon_mathfun.h`](https://github.com/Tencent/ncnn/blob/main/src/layer/arm/neon_mathfun.h).
- On ARM processors—particularly Cortex-A53 and A55—denormals trigger slow micro-code paths that severely degrade NEON SIMD performance, making aggressive flushing essential for real-time inference.
- The library auto-configures `flush_denormals = 3` for A53/A55 cores in [`src/option.cpp`](https://github.com/Tencent/ncnn/blob/main/src/option.cpp), ensuring optimal mobile performance by default.

## Frequently Asked Questions

### What are denormal floating-point values and why do they cause slowdowns?

Denormal (or subnormal) floating-point values are non-zero numbers smaller than the normal representable range, used to handle gradual underflow. They cause slowdowns because many processors—including ARM Cortex-A53 and A55—lack hardware optimization for these edge cases, forcing execution into slow micro-code routines that can take 10-100x longer than normal floating-point operations.

### How do I know if my ARM device is affected by denormal performance penalties?

If your device uses a Cortex-A53, Cortex-A55, or similar entry-level to mid-range ARM core, it is likely affected. ncnn automatically detects these architectures in [`src/option.cpp`](https://github.com/Tencent/ncnn/blob/main/src/option.cpp) and sets `flush_denormals = 3` by default. You can verify the setting by checking `net.opt.flush_denormals` after creating the Net object, or profile inference with the option set to 0 versus 3 to measure the performance difference.

### Can disabling flush_denormals improve numerical accuracy?

Disabling `flush_denormals` (setting it to 0) preserves denormal values rather than flushing them to zero, which theoretically maintains higher precision for very small numbers. However, in deep learning inference, the impact on accuracy is typically negligible because neural network weights and activations are usually well-scaled, and the performance cost of handling denormals on ARM processors far outweighs any minimal precision benefits.

### Where does ncnn handle the actual state switching for flush_denormals?

The state management occurs in [`src/net.cpp`](https://github.com/Tencent/ncnn/blob/main/src/net.cpp) around lines 2441-2443, where the code calls `get_flush_denormals` to save the current CPU state, applies the new mode for the inference run, and restores the original state afterward. On x86, the actual hardware control is implemented in [`src/cpu.cpp`](https://github.com/Tencent/ncnn/blob/main/src/cpu.cpp) using SSE intrinsics, while ARM kernels handle it explicitly in files like [`src/layer/arm/neon_mathfun.h`](https://github.com/Tencent/ncnn/blob/main/src/layer/arm/neon_mathfun.h) through NEON clamping operations.