bf16_storage vs fp16_storage in ncnn: Format Differences, Hardware Support, and Selection Guide
The primary distinction between bf16_storage and fp16_storage in ncnn lies in their numeric representation: fp16 uses IEEE-754 binary16 (5-bit exponent, 10-bit mantissa) for higher precision but limited range, while bf16 uses brain-float16 (8-bit exponent, 7-bit mantissa) for a much wider dynamic range at the cost of coarser precision; choose fp16 for vision models with native half-precision support, and bf16 for transformers or wide-range activations that risk FP16 overflow.
ncnn supports three numeric formats for intermediate tensors: FP32 (full precision), FP16 (16-bit IEEE-754), and BF16 (16-bit brain-float16). Both bf16_storage and fp16_storage reduce memory bandwidth and cache pressure, but they exhibit fundamentally different numerical behavior and hardware compatibility. Understanding these differences is critical for optimizing inference on mobile CPUs, GPUs, and specialized accelerators.
Numeric Characteristics and Precision Trade-offs
The structural difference between the two 16-bit formats determines their numerical behavior in deep learning workloads.
| Feature | FP16 (binary16) | BF16 (brain-float16) |
|---|---|---|
| Exponent bits | 5 (bias = 15) | 8 (bias = 127) |
| Mantissa bits | 10 (~3 decimal digits) | 7 (~2 decimal digits) |
| Approximate range | 6 × 10⁻⁸ to 65,504 | 1 × 10⁻³⁸ to 3.4 × 10³⁸ |
FP16 provides higher precision for small values but saturates quickly for large activations, making it prone to overflow in deep residual networks or attention mechanisms. BF16 retains the same exponent range as FP32, preventing overflow in models with wide dynamic ranges, but introduces higher relative error due to its truncated mantissa.
In src/option.h, these are controlled by separate flags:
// src/option.h
bool use_fp16_storage; // Line ~92
bool use_bf16_storage; // Line ~86
Hardware Support and Runtime Detection
ncnn gates storage format selection behind runtime CPU capability checks and compile-time macros. The framework evaluates hardware support before converting tensors in src/net.cpp.
ARM Architectures:
- ARMv8.2 ASIMD-HP: Native FP16 arithmetic and storage (
NCNN_ARM82,cpu_support_arm_asimdhp()) - ARMv8.2 VFPv4: FP16 conversion instructions; supports BF16 storage with FP16 arithmetic fallback (
NCNN_VFPV4,cpu_support_arm_vfpv4())
x86 Architectures:
- AVX-512 BF16: Direct BF16 ISA support for both storage and specialized instructions (
NCNN_AVX512,cpu_support_x86_avx512_bf16()) - AVX-512 FP16: Native FP16 support for compatible Intel/AMD processors
RISC-V:
- Zfh/Zvfh extensions: FP16 support with BF16 storage capability (
NCNN_ZFH)
GPU (Vulkan):
- Runtime detection via
vkdev->info.support_bf16_storage()for BF16 buffers
The conversion logic in src/net.cpp (lines 69-100) prioritizes hardware capabilities:
// FP16 path (src/net.cpp)
#if NCNN_ARM82
if (opt.use_fp16_storage && cpu_support_arm_asimdhp() && layer->support_fp16_storage) {
cast_float32_to_float16(...);
}
#endif
// BF16 path (src/net.cpp)
#if NCNN_BF16
if (opt.use_bf16_storage && layer->support_bf16_storage) {
cast_float32_to_bfloat16(...);
}
#endif
When to Use FP16 Storage
Enable use_fp16_storage when your deployment target supports native half-precision arithmetic and your model's activation range fits within FP16's limited scope.
Optimal scenarios:
- Mobile vision models: Networks like MobileNet-V2, EfficientNet-B0, and ResNet-18 that have been trained or fine-tuned with FP16-aware techniques
- ARM ASIMD-HP devices: Cortex-A75, A76, and newerbig cores that provide native FP16 arithmetic (
use_fp16_arithmetic = true) - Normal dynamic ranges: When intermediate activations consistently remain below 65,504 and above 6 × 10⁻⁸
Configuration example:
ncnn::Option opt;
opt.use_fp16_storage = true; // Enable 16-bit blob storage
opt.use_fp16_arithmetic = true; // Compute in FP16 where possible
net.opt = opt;
When to Use BF16 Storage
Prefer use_bf16_storage for models exhibiting large activation magnitudes or when deploying on hardware that lacks native FP16 arithmetic but efficiently handles BF16 conversion.
Optimal scenarios:
- Transformer architectures: BERT, GPT-style models, and attention mechanisms with wide dynamic ranges where FP16 would overflow
- VFPv4-only devices: ARMv8.2 cores without ASIMD-HP that can convert FP32↔BF16 efficiently while computing in FP16 or FP32
- Vulkan GPU backends: When
vkdev->info.support_bf16_storage()returns true, enabling BF16 buffers regardless of arithmetic precision
Note that use_bf16_storage shares the arithmetic flag with FP16: use_fp16_arithmetic. When BF16 storage is enabled on VFPv4 or similar, ncnn performs arithmetic in FP16 (or FP32) but maintains BF16 tensor storage to reduce memory footprint.
Configuration example:
ncnn::Option opt;
opt.use_bf16_storage = true; // Keep tensors in BF16
opt.use_fp16_arithmetic = true; // Use FP16 math if CPU supports it
net.opt = opt;
Interaction with Packing Layout
Both storage flags integrate with ncnn's packing optimizations. When use_packing_layout is enabled, tensors are stored as packed vectors (e.g., pack4 for NEON). The storage flags determine the element size of these packed blobs—either 16-bit FP16 or 16-bit BF16—while maintaining the packed memory layout for vectorized operations.
Layer implementations in src/layer/arm/ and src/layer/x86/ check opt.use_bf16_storage or opt.use_fp16_storage to select appropriate SIMD kernels. For example, convolution layers invoke specific BF16 cast routines in src/layer/x86/cast_bf16.h when AVX-512 BF16 is available.
Configuration Examples by Platform
ARM Cortex-A75 (Mobile Vision):
ncnn::Net net;
ncnn::Option opt;
opt.use_fp16_storage = true;
opt.use_fp16_arithmetic = true;
net.opt = opt;
net.load_param("mobilenetv2.param");
ARM Cortex-A55 with VFPv4 (Transformer):
ncnn::Net net;
ncnn::Option opt;
opt.use_bf16_storage = true; // Prevent overflow in attention layers
opt.use_fp16_arithmetic = true;
net.opt = opt;
net.load_param("bert.param");
Vulkan GPU with Runtime Detection:
ncnn::VulkanDevice* vkdev = ncnn::get_default_vkdev();
ncnn::Option opt;
opt.use_vulkan_compute = true;
opt.use_bf16_storage = vkdev->info.support_bf16_storage();
net.opt = opt;
net.set_vulkan_device(vkdev);
Summary
- FP16 (binary16) offers higher precision for small values but risks overflow with large activations; ideal for standard vision models on modern ARM ASIMD-HP or AVX-512 FP16 hardware.
- BF16 (brain-float16) retains FP32's dynamic range with reduced precision; preferable for transformers and wide-range activations, and functions efficiently on VFPv4 and AVX-512 BF16 hardware.
- Storage vs. Arithmetic: Both formats use 16-bit storage, but BF16 often relies on
use_fp16_arithmeticfor computation, while FP16 may use native FP16 math. - Hardware Detection: ncnn automatically checks CPU capabilities via
cpu_support_arm_asimdhp(),cpu_support_x86_avx512_bf16(), and Vulkan device properties before enabling conversions. - Source References: Configuration resides in
src/option.h, conversion logic insrc/net.cpp, and architecture-specific implementations insrc/layer/arm/andsrc/layer/x86/cast_bf16.h.
Frequently Asked Questions
What is the main numerical difference between bf16 and fp16 in ncnn?
FP16 allocates 5 bits to the exponent and 10 bits to the mantissa, providing approximately 3 decimal digits of precision but saturating at 65,504. BF16 uses 8 exponent bits and 7 mantissa bits, matching FP32's massive dynamic range (~10⁻³⁸ to 10³⁸) but with only 2 decimal digits of precision. This makes BF16 resistant to overflow in deep networks with large activations, while FP16 preserves more detail for small-value features.
Can I enable both bf16_storage and fp16_storage simultaneously in ncnn?
No, ncnn treats these as mutually exclusive storage options. In src/net.cpp, the framework checks use_bf16_storage before use_fp16_storage, and enabling BF16 storage typically bypasses FP16 storage paths. However, BF16 storage can be combined with FP16 arithmetic via use_fp16_arithmetic = true, allowing computations to use half-precision math while tensors reside in BF16 format.
Which mobile processors support bf16_storage in ncnn?
Most ARMv8.2-A and newer processors support BF16 storage, including Cortex-A55, A75, A76, and X-series cores. Specific support depends on the presence of VFPv4 or ASIMD-HP instructions. On x86, BF16 storage requires AVX-512 BF16 support (found in Intel Cooper Lake, Sapphire Rapids, and newer AMD EPYC processors). The cpu_support_arm_vfpv4() and cpu_support_x86_avx512_bf16() functions in ncnn's CPU detection layer determine availability at runtime.
How does bf16_storage affect model accuracy compared to fp16_storage?
BF16 generally exhibits higher relative error for small-magnitude values due to its 7-bit mantissa versus FP16's 10-bit mantissa. However, BF16 rarely overflows, preventing the catastrophic accuracy loss that occurs when FP16 saturates. For vision models with normalized inputs (0-1 range), FP16 often preserves accuracy better; for language models or deep residual networks with unconstrained activations, BF16 typically delivers more stable inference by avoiding infinity values.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →