ncnn Supported Data Types and Type Conversion During Inference

ncnn supports float32, float16, bfloat16, and int8 data types, managing conversions through explicit cast helpers, quantization functions, and runtime Option flags that control automatic type conversion during inference.

The Tencent/ncnn inference framework provides flexible precision management for deploying deep learning models on mobile and edge devices. Understanding how ncnn handles data type conversions allows developers to optimize memory usage and computational performance without modifying model files.

Overview of ncnn Supported Data Types

ncnn represents all tensor data through the generic ncnn::Mat container, which tracks element size via the elemsize property. The framework natively supports four primary data formats:

float32 (Default Precision)

float32 serves as the default computation format, consuming 4 bytes per element (elemsize == 4). This high-precision mode ensures maximum accuracy but requires the most memory bandwidth. When no optimization flags are set, ncnn maintains all activations and weights in float32 throughout inference.

float16 (Half-Precision)

float16 reduces storage to 2 bytes per element, enabling memory-efficient inference when Option::use_fp16_* flags are enabled. This format is particularly beneficial for mobile GPUs and ARM processors with NEON FP16 support. The conversion between float32 and float16 utilizes SIMD-accelerated routines in architecture-specific headers like src/layer/arm/cast_fp16.h and src/layer/x86/cast_fp16.h.

bfloat16

bfloat16 also uses 2 bytes but employs a different exponent format optimized for deep learning workloads. This type activates on CPUs supporting AVX-512 BF16 extensions, providing a middle ground between float32 accuracy and float16 storage efficiency. Conversion functions float32_to_bfloat16 and bfloat16_to_float32 handle the element-wise transformation.

int8 (Quantized)

int8 quantization compresses data to 1 byte per element, enabling integer-only inference on supported hardware. This mode requires per-channel scale factors calculated during model calibration. The quantize_to_int8 function converts float32 activations to int8 using these scales, while cast_int8_to_float32 handles dequantization.

How ncnn Manages Type Conversions During Inference

ncnn implements a three-tiered conversion strategy: explicit cast helpers for manual control, quantization functions for INT8 deployment, and automatic conversion via runtime configuration flags.

Explicit Cast Helpers

The framework provides dedicated conversion functions in src/mat.h for direct data transformation:

// Convert between float32 and float16
void cast_float32_to_float16(const Mat& src, Mat& dst, const Option& opt = Option());
void cast_float16_to_float32(const Mat& src, Mat& dst, const Option& opt = Option());

// Convert between float32 and bfloat16
void cast_float32_to_bfloat16(const Mat& src, Mat& dst, const Option& opt = Option());
void bfloat16_to_float32(const Mat& src, Mat& dst, const Option& opt = Option());

// Convert between int8 and float32
void cast_int8_to_float32(const Mat& src, Mat& dst, const Option& opt = Option());

These functions perform element-wise conversion using optimized SIMD intrinsics. For example, x86 implementations leverage AVX-512 for vectorized float16 conversion, while ARM builds utilize NEON FP16 instructions.

Quantization to int8

For quantized inference, ncnn uses the quantize_to_int8 function declared at line 95-96 of src/mat.h:

void quantize_to_int8(const Mat& src, Mat& dst, const Mat& scale_data, const Option& opt = Option());

The implementation in src/mat.cpp (line 1719) multiplies float32 elements by the reciprocal of their per-channel scale and rounds to the nearest integer. This requires pre-calculated scale tensors typically generated during model conversion. Layers like InnerProduct demonstrate this pattern at line 73-74 of src/layer/innerproduct.cpp, where weight quantization occurs before inference.

Automatic Conversion via Option Flags

ncnn automates type management through the Option struct in src/option.h, allowing runtime precision control without code modification:

bool use_fp16_packed;      // Pack 4/8 elements into fp16 registers
bool use_fp16_storage;     // Store tensors in fp16 memory
bool use_fp16_arithmetic;  // Perform arithmetic in fp16 when possible
bool use_fp16_uniform;     // Upload uniform constants as fp16

During inference, layers check these flags to determine conversion strategy. For example, src/layer/convolution.cpp at line 102-103 implements logic to cast input blobs to fp16 when use_fp16_arithmetic is enabled. The typical execution flow follows this pattern:

  1. Weight Loading: When use_fp16_storage is true, layers call cast_float32_to_float16() during model initialization to compress weight blobs.
  2. Activation Processing: Before heavy computation, inputs may be cast to fp16 if use_fp16_arithmetic is set, with results cast back to float32 for subsequent layers.
  3. Quantized Execution: INT8 models invoke quantize_to_int8() on activations using pre-computed scales stored alongside model weights.

Practical Code Examples

Converting float32 to float16

This example demonstrates explicit conversion between precision formats using ncnn's cast helpers:

#include "mat.h"
#include "option.h"

int main() {
    // Create float32 tensor (10 elements)
    ncnn::Mat src(10, 4u);  // 4 bytes = float32
    src.fill(1.23f);
    
    // Prepare destination for float16
    ncnn::Mat dst;
    ncnn::Option opt;
    opt.use_fp16_storage = true;
    
    // Perform conversion (dst.elemsize becomes 2)
    ncnn::cast_float32_to_float16(src, dst, opt);
    
    return 0;
}

Quantizing Activations to int8

For quantized inference, convert floating-point activations using per-channel scales:

#include "mat.h"
#include "option.h"

int main() {
    // NCHW float32 activation (batch=1, channels=3, height=224, width=224)
    ncnn::Mat activation(1, 3, 224, 224, 4u);
    
    // Per-channel scale factors from calibration
    ncnn::Mat scale_data(1, 3, 1, 1, 4u);
    // ... populate scale_data ...
    
    ncnn::Mat activation_int8;
    ncnn::Option opt;
    
    // Quantize to int8 (elemsize becomes 1)
    ncnn::quantize_to_int8(activation, activation_int8, scale_data, opt);
    
    return 0;
}

Configuring fp16 Arithmetic in Inference

Enable half-precision computation through runtime options without modifying model files:

ncnn::Option opt;
opt.use_fp16_arithmetic = true;  // Compute in fp16 where supported
opt.use_fp16_storage = true;     // Store tensors in fp16

ncnn::Net net;
net.load_param("model.param");
net.load_model("model.bin");
net.opt = opt;

ncnn::Mat in, out;
// ... prepare input ...
net.extract("output", out);

Summary

  • ncnn supports four primary data types: float32 (4 bytes), float16 (2 bytes), bfloat16 (2 bytes), and int8 (1 byte), all managed through the ncnn::Mat container using the elemsize property.
  • Explicit conversion functions in src/mat.h provide direct control over type casting, including cast_float32_to_float16, cast_float32_to_bfloat16, and cast_int8_to_float32, implemented with SIMD optimizations for x86 and ARM architectures.
  • Quantization workflow uses quantize_to_int8 with per-channel scale factors to convert float32 activations to int8, enabling integer-only inference on supported hardware.
  • Runtime precision control via Option flags (use_fp16_storage, use_fp16_arithmetic, use_fp16_packed) allows automatic type conversion during inference without model modification, with layers in src/layer/convolution.cpp and src/layer/innerproduct.cpp implementing the conversion logic.

Frequently Asked Questions

What data types does ncnn support for model inference?

ncnn supports float32 (32-bit floating point), float16 (16-bit half-precision), bfloat16 (16-bit brain floating point), and int8 (8-bit quantized integer). These types are stored in the ncnn::Mat container, which tracks the element size via elemsize (4 bytes for float32, 2 bytes for float16/bfloat16, and 1 byte for int8).

How does ncnn convert between float32 and float16 during inference?

ncnn provides explicit conversion functions declared in src/mat.h: cast_float32_to_float16 and cast_float16_to_float32. These functions perform element-wise conversion using SIMD intrinsics optimized for the target architecture (NEON on ARM, AVX-512 on x86). During automatic inference, layers check Option::use_fp16_storage and Option::use_fp16_arithmetic to determine whether to cast weights and activations to float16 before computation.

What is the difference between float16 and bfloat16 in ncnn?

float16 (half-precision) uses 16 bits with a 5-bit exponent and 10-bit mantissa, providing higher precision but narrower range, ideal for memory-constrained mobile GPUs. bfloat16 (brain floating point) uses 16 bits with an 8-bit exponent and 7-bit mantissa, matching float32's range but with reduced precision, optimized for AVX-512 CPUs. ncnn supports both via cast_float32_to_float16/cast_float16_to_float32 and float32_to_bfloat16/bfloat16_to_float32 respectively.

How does ncnn handle int8 quantization for model weights and activations?

ncnn implements int8 quantization through the quantize_to_int8 function in src/mat.cpp (line 1719), which converts float32 tensors to int8 using per-channel scale factors. During inference, layers like InnerProduct (line 73-74) and Convolution quantize weights during loading if int8 mode is enabled, while activations are quantized on-the-fly using pre-calculated scales stored in scale_data tensors. The cast_int8_to_float32 function handles dequantization when layers require floating-point input.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →