# bf16_storage vs fp16_storage in ncnn: Format Differences, Hardware Support, and Selection Guide

> Explore bf16_storage vs fp16_storage in ncnn. Understand their format differences, hardware support, and get a guide to choosing the right one for your models, optimizing performance and precision.

- Repository: [Tencent/ncnn](https://github.com/tencent/ncnn)
- Tags: deep-dive
- Published: 2026-02-23

---

**The primary distinction between `bf16_storage` and `fp16_storage` in ncnn lies in their numeric representation: fp16 uses IEEE-754 binary16 (5-bit exponent, 10-bit mantissa) for higher precision but limited range, while bf16 uses brain-float16 (8-bit exponent, 7-bit mantissa) for a much wider dynamic range at the cost of coarser precision; choose fp16 for vision models with native half-precision support, and bf16 for transformers or wide-range activations that risk FP16 overflow.**

ncnn supports three numeric formats for intermediate tensors: **FP32** (full precision), **FP16** (16-bit IEEE-754), and **BF16** (16-bit brain-float16). Both `bf16_storage` and `fp16_storage` reduce memory bandwidth and cache pressure, but they exhibit fundamentally different numerical behavior and hardware compatibility. Understanding these differences is critical for optimizing inference on mobile CPUs, GPUs, and specialized accelerators.

## Numeric Characteristics and Precision Trade-offs

The structural difference between the two 16-bit formats determines their numerical behavior in deep learning workloads.

| Feature | FP16 (binary16) | BF16 (brain-float16) |
|---------|----------------|---------------------|
| **Exponent bits** | 5 (bias = 15) | 8 (bias = 127) |
| **Mantissa bits** | 10 (~3 decimal digits) | 7 (~2 decimal digits) |
| **Approximate range** | 6 × 10⁻⁸ to 65,504 | 1 × 10⁻³⁸ to 3.4 × 10³⁸ |

**FP16** provides higher precision for small values but saturates quickly for large activations, making it prone to overflow in deep residual networks or attention mechanisms. **BF16** retains the same exponent range as FP32, preventing overflow in models with wide dynamic ranges, but introduces higher relative error due to its truncated mantissa.

In [`src/option.h`](https://github.com/Tencent/ncnn/blob/main/src/option.h), these are controlled by separate flags:

```cpp
// src/option.h
bool use_fp16_storage;  // Line ~92
bool use_bf16_storage;  // Line ~86

```

## Hardware Support and Runtime Detection

ncnn gates storage format selection behind runtime CPU capability checks and compile-time macros. The framework evaluates hardware support before converting tensors in [`src/net.cpp`](https://github.com/Tencent/ncnn/blob/main/src/net.cpp).

**ARM Architectures:**
- **ARMv8.2 ASIMD-HP**: Native FP16 arithmetic and storage (`NCNN_ARM82`, `cpu_support_arm_asimdhp()`)
- **ARMv8.2 VFPv4**: FP16 conversion instructions; supports BF16 storage with FP16 arithmetic fallback (`NCNN_VFPV4`, `cpu_support_arm_vfpv4()`)

**x86 Architectures:**
- **AVX-512 BF16**: Direct BF16 ISA support for both storage and specialized instructions (`NCNN_AVX512`, `cpu_support_x86_avx512_bf16()`)
- **AVX-512 FP16**: Native FP16 support for compatible Intel/AMD processors

**RISC-V:**
- **Zfh/Zvfh extensions**: FP16 support with BF16 storage capability (`NCNN_ZFH`)

**GPU (Vulkan):**
- Runtime detection via `vkdev->info.support_bf16_storage()` for BF16 buffers

The conversion logic in [`src/net.cpp`](https://github.com/Tencent/ncnn/blob/main/src/net.cpp) (lines 69-100) prioritizes hardware capabilities:

```cpp
// FP16 path (src/net.cpp)
#if NCNN_ARM82
    if (opt.use_fp16_storage && cpu_support_arm_asimdhp() && layer->support_fp16_storage) {
        cast_float32_to_float16(...);
    }
#endif

// BF16 path (src/net.cpp)  
#if NCNN_BF16
    if (opt.use_bf16_storage && layer->support_bf16_storage) {
        cast_float32_to_bfloat16(...);
    }
#endif

```

## When to Use FP16 Storage

Enable `use_fp16_storage` when your deployment target supports native half-precision arithmetic and your model's activation range fits within FP16's limited scope.

**Optimal scenarios:**
- **Mobile vision models**: Networks like MobileNet-V2, EfficientNet-B0, and ResNet-18 that have been trained or fine-tuned with FP16-aware techniques
- **ARM ASIMD-HP devices**: Cortex-A75, A76, and newerbig cores that provide native FP16 arithmetic (`use_fp16_arithmetic = true`)
- **Normal dynamic ranges**: When intermediate activations consistently remain below 65,504 and above 6 × 10⁻⁸

**Configuration example:**

```cpp
ncnn::Option opt;
opt.use_fp16_storage = true;      // Enable 16-bit blob storage
opt.use_fp16_arithmetic = true;   // Compute in FP16 where possible
net.opt = opt;

```

## When to Use BF16 Storage

Prefer `use_bf16_storage` for models exhibiting large activation magnitudes or when deploying on hardware that lacks native FP16 arithmetic but efficiently handles BF16 conversion.

**Optimal scenarios:**
- **Transformer architectures**: BERT, GPT-style models, and attention mechanisms with wide dynamic ranges where FP16 would overflow
- **VFPv4-only devices**: ARMv8.2 cores without ASIMD-HP that can convert FP32↔BF16 efficiently while computing in FP16 or FP32
- **Vulkan GPU backends**: When `vkdev->info.support_bf16_storage()` returns true, enabling BF16 buffers regardless of arithmetic precision

Note that `use_bf16_storage` shares the arithmetic flag with FP16: `use_fp16_arithmetic`. When BF16 storage is enabled on VFPv4 or similar, ncnn performs arithmetic in FP16 (or FP32) but maintains BF16 tensor storage to reduce memory footprint.

**Configuration example:**

```cpp
ncnn::Option opt;
opt.use_bf16_storage = true;      // Keep tensors in BF16
opt.use_fp16_arithmetic = true;   // Use FP16 math if CPU supports it
net.opt = opt;

```

## Interaction with Packing Layout

Both storage flags integrate with ncnn's packing optimizations. When `use_packing_layout` is enabled, tensors are stored as packed vectors (e.g., `pack4` for NEON). The storage flags determine the element size of these packed blobs—either 16-bit FP16 or 16-bit BF16—while maintaining the packed memory layout for vectorized operations.

Layer implementations in `src/layer/arm/` and `src/layer/x86/` check `opt.use_bf16_storage` or `opt.use_fp16_storage` to select appropriate SIMD kernels. For example, convolution layers invoke specific BF16 cast routines in [`src/layer/x86/cast_bf16.h`](https://github.com/Tencent/ncnn/blob/main/src/layer/x86/cast_bf16.h) when AVX-512 BF16 is available.

## Configuration Examples by Platform

**ARM Cortex-A75 (Mobile Vision):**

```cpp
ncnn::Net net;
ncnn::Option opt;
opt.use_fp16_storage = true;
opt.use_fp16_arithmetic = true;
net.opt = opt;
net.load_param("mobilenetv2.param");

```

**ARM Cortex-A55 with VFPv4 (Transformer):**

```cpp
ncnn::Net net;
ncnn::Option opt;
opt.use_bf16_storage = true;  // Prevent overflow in attention layers
opt.use_fp16_arithmetic = true;
net.opt = opt;
net.load_param("bert.param");

```

**Vulkan GPU with Runtime Detection:**

```cpp
ncnn::VulkanDevice* vkdev = ncnn::get_default_vkdev();
ncnn::Option opt;
opt.use_vulkan_compute = true;
opt.use_bf16_storage = vkdev->info.support_bf16_storage();
net.opt = opt;
net.set_vulkan_device(vkdev);

```

## Summary

- **FP16 (binary16)** offers higher precision for small values but risks overflow with large activations; ideal for standard vision models on modern ARM ASIMD-HP or AVX-512 FP16 hardware.
- **BF16 (brain-float16)** retains FP32's dynamic range with reduced precision; preferable for transformers and wide-range activations, and functions efficiently on VFPv4 and AVX-512 BF16 hardware.
- **Storage vs. Arithmetic**: Both formats use 16-bit storage, but BF16 often relies on `use_fp16_arithmetic` for computation, while FP16 may use native FP16 math.
- **Hardware Detection**: ncnn automatically checks CPU capabilities via `cpu_support_arm_asimdhp()`, `cpu_support_x86_avx512_bf16()`, and Vulkan device properties before enabling conversions.
- **Source References**: Configuration resides in [`src/option.h`](https://github.com/Tencent/ncnn/blob/main/src/option.h), conversion logic in [`src/net.cpp`](https://github.com/Tencent/ncnn/blob/main/src/net.cpp), and architecture-specific implementations in `src/layer/arm/` and [`src/layer/x86/cast_bf16.h`](https://github.com/Tencent/ncnn/blob/main/src/layer/x86/cast_bf16.h).

## Frequently Asked Questions

### What is the main numerical difference between bf16 and fp16 in ncnn?

FP16 allocates 5 bits to the exponent and 10 bits to the mantissa, providing approximately 3 decimal digits of precision but saturating at 65,504. BF16 uses 8 exponent bits and 7 mantissa bits, matching FP32's massive dynamic range (~10⁻³⁸ to 10³⁸) but with only 2 decimal digits of precision. This makes BF16 resistant to overflow in deep networks with large activations, while FP16 preserves more detail for small-value features.

### Can I enable both bf16_storage and fp16_storage simultaneously in ncnn?

No, ncnn treats these as mutually exclusive storage options. In [`src/net.cpp`](https://github.com/Tencent/ncnn/blob/main/src/net.cpp), the framework checks `use_bf16_storage` before `use_fp16_storage`, and enabling BF16 storage typically bypasses FP16 storage paths. However, BF16 storage can be combined with FP16 arithmetic via `use_fp16_arithmetic = true`, allowing computations to use half-precision math while tensors reside in BF16 format.

### Which mobile processors support bf16_storage in ncnn?

Most ARMv8.2-A and newer processors support BF16 storage, including Cortex-A55, A75, A76, and X-series cores. Specific support depends on the presence of VFPv4 or ASIMD-HP instructions. On x86, BF16 storage requires AVX-512 BF16 support (found in Intel Cooper Lake, Sapphire Rapids, and newer AMD EPYC processors). The `cpu_support_arm_vfpv4()` and `cpu_support_x86_avx512_bf16()` functions in ncnn's CPU detection layer determine availability at runtime.

### How does bf16_storage affect model accuracy compared to fp16_storage?

BF16 generally exhibits higher relative error for small-magnitude values due to its 7-bit mantissa versus FP16's 10-bit mantissa. However, BF16 rarely overflows, preventing the catastrophic accuracy loss that occurs when FP16 saturates. For vision models with normalized inputs (0-1 range), FP16 often preserves accuracy better; for language models or deep residual networks with unconstrained activations, BF16 typically delivers more stable inference by avoiding infinity values.