# ncnn Supported Data Types and Type Conversion During Inference

> Explore ncnn supported data types float fp16 bf16 int8 and learn how ncnn handles type conversions during inference with explicit casts quantization and runtime options.

- Repository: [Tencent/ncnn](https://github.com/tencent/ncnn)
- Tags: internals
- Published: 2026-02-23

---

**ncnn supports float32, float16, bfloat16, and int8 data types, managing conversions through explicit cast helpers, quantization functions, and runtime Option flags that control automatic type conversion during inference.**

The Tencent/ncnn inference framework provides flexible precision management for deploying deep learning models on mobile and edge devices. Understanding how ncnn handles data type conversions allows developers to optimize memory usage and computational performance without modifying model files.

## Overview of ncnn Supported Data Types

ncnn represents all tensor data through the generic `ncnn::Mat` container, which tracks element size via the `elemsize` property. The framework natively supports four primary data formats:

### float32 (Default Precision)

**float32** serves as the default computation format, consuming 4 bytes per element (`elemsize == 4`). This high-precision mode ensures maximum accuracy but requires the most memory bandwidth. When no optimization flags are set, ncnn maintains all activations and weights in float32 throughout inference.

### float16 (Half-Precision)

**float16** reduces storage to 2 bytes per element, enabling memory-efficient inference when `Option::use_fp16_*` flags are enabled. This format is particularly beneficial for mobile GPUs and ARM processors with NEON FP16 support. The conversion between float32 and float16 utilizes SIMD-accelerated routines in architecture-specific headers like [`src/layer/arm/cast_fp16.h`](https://github.com/Tencent/ncnn/blob/main/src/layer/arm/cast_fp16.h) and [`src/layer/x86/cast_fp16.h`](https://github.com/Tencent/ncnn/blob/main/src/layer/x86/cast_fp16.h).

### bfloat16

**bfloat16** also uses 2 bytes but employs a different exponent format optimized for deep learning workloads. This type activates on CPUs supporting AVX-512 BF16 extensions, providing a middle ground between float32 accuracy and float16 storage efficiency. Conversion functions `float32_to_bfloat16` and `bfloat16_to_float32` handle the element-wise transformation.

### int8 (Quantized)

**int8** quantization compresses data to 1 byte per element, enabling integer-only inference on supported hardware. This mode requires per-channel scale factors calculated during model calibration. The `quantize_to_int8` function converts float32 activations to int8 using these scales, while `cast_int8_to_float32` handles dequantization.

## How ncnn Manages Type Conversions During Inference

ncnn implements a three-tiered conversion strategy: explicit cast helpers for manual control, quantization functions for INT8 deployment, and automatic conversion via runtime configuration flags.

### Explicit Cast Helpers

The framework provides dedicated conversion functions in [`src/mat.h`](https://github.com/Tencent/ncnn/blob/main/src/mat.h) for direct data transformation:

```cpp
// Convert between float32 and float16
void cast_float32_to_float16(const Mat& src, Mat& dst, const Option& opt = Option());
void cast_float16_to_float32(const Mat& src, Mat& dst, const Option& opt = Option());

// Convert between float32 and bfloat16
void cast_float32_to_bfloat16(const Mat& src, Mat& dst, const Option& opt = Option());
void bfloat16_to_float32(const Mat& src, Mat& dst, const Option& opt = Option());

// Convert between int8 and float32
void cast_int8_to_float32(const Mat& src, Mat& dst, const Option& opt = Option());

```

These functions perform element-wise conversion using optimized SIMD intrinsics. For example, x86 implementations leverage AVX-512 for vectorized float16 conversion, while ARM builds utilize NEON FP16 instructions.

### Quantization to int8

For quantized inference, ncnn uses the `quantize_to_int8` function declared at line 95-96 of [`src/mat.h`](https://github.com/Tencent/ncnn/blob/main/src/mat.h):

```cpp
void quantize_to_int8(const Mat& src, Mat& dst, const Mat& scale_data, const Option& opt = Option());

```

The implementation in [`src/mat.cpp`](https://github.com/Tencent/ncnn/blob/main/src/mat.cpp) (line 1719) multiplies float32 elements by the reciprocal of their per-channel scale and rounds to the nearest integer. This requires pre-calculated scale tensors typically generated during model conversion. Layers like `InnerProduct` demonstrate this pattern at line 73-74 of [`src/layer/innerproduct.cpp`](https://github.com/Tencent/ncnn/blob/main/src/layer/innerproduct.cpp), where weight quantization occurs before inference.

### Automatic Conversion via Option Flags

ncnn automates type management through the `Option` struct in [`src/option.h`](https://github.com/Tencent/ncnn/blob/main/src/option.h), allowing runtime precision control without code modification:

```cpp
bool use_fp16_packed;      // Pack 4/8 elements into fp16 registers
bool use_fp16_storage;     // Store tensors in fp16 memory
bool use_fp16_arithmetic;  // Perform arithmetic in fp16 when possible
bool use_fp16_uniform;     // Upload uniform constants as fp16

```

During inference, layers check these flags to determine conversion strategy. For example, [`src/layer/convolution.cpp`](https://github.com/Tencent/ncnn/blob/main/src/layer/convolution.cpp) at line 102-103 implements logic to cast input blobs to fp16 when `use_fp16_arithmetic` is enabled. The typical execution flow follows this pattern:

1. **Weight Loading**: When `use_fp16_storage` is true, layers call `cast_float32_to_float16()` during model initialization to compress weight blobs.
2. **Activation Processing**: Before heavy computation, inputs may be cast to fp16 if `use_fp16_arithmetic` is set, with results cast back to float32 for subsequent layers.
3. **Quantized Execution**: INT8 models invoke `quantize_to_int8()` on activations using pre-computed scales stored alongside model weights.

## Practical Code Examples

### Converting float32 to float16

This example demonstrates explicit conversion between precision formats using ncnn's cast helpers:

```cpp
#include "mat.h"
#include "option.h"

int main() {
    // Create float32 tensor (10 elements)
    ncnn::Mat src(10, 4u);  // 4 bytes = float32
    src.fill(1.23f);
    
    // Prepare destination for float16
    ncnn::Mat dst;
    ncnn::Option opt;
    opt.use_fp16_storage = true;
    
    // Perform conversion (dst.elemsize becomes 2)
    ncnn::cast_float32_to_float16(src, dst, opt);
    
    return 0;
}

```

### Quantizing Activations to int8

For quantized inference, convert floating-point activations using per-channel scales:

```cpp
#include "mat.h"
#include "option.h"

int main() {
    // NCHW float32 activation (batch=1, channels=3, height=224, width=224)
    ncnn::Mat activation(1, 3, 224, 224, 4u);
    
    // Per-channel scale factors from calibration
    ncnn::Mat scale_data(1, 3, 1, 1, 4u);
    // ... populate scale_data ...
    
    ncnn::Mat activation_int8;
    ncnn::Option opt;
    
    // Quantize to int8 (elemsize becomes 1)
    ncnn::quantize_to_int8(activation, activation_int8, scale_data, opt);
    
    return 0;
}

```

### Configuring fp16 Arithmetic in Inference

Enable half-precision computation through runtime options without modifying model files:

```cpp
ncnn::Option opt;
opt.use_fp16_arithmetic = true;  // Compute in fp16 where supported
opt.use_fp16_storage = true;     // Store tensors in fp16

ncnn::Net net;
net.load_param("model.param");
net.load_model("model.bin");
net.opt = opt;

ncnn::Mat in, out;
// ... prepare input ...
net.extract("output", out);

```

## Summary

- **ncnn supports four primary data types**: float32 (4 bytes), float16 (2 bytes), bfloat16 (2 bytes), and int8 (1 byte), all managed through the `ncnn::Mat` container using the `elemsize` property.
- **Explicit conversion functions** in [`src/mat.h`](https://github.com/Tencent/ncnn/blob/main/src/mat.h) provide direct control over type casting, including `cast_float32_to_float16`, `cast_float32_to_bfloat16`, and `cast_int8_to_float32`, implemented with SIMD optimizations for x86 and ARM architectures.
- **Quantization workflow** uses `quantize_to_int8` with per-channel scale factors to convert float32 activations to int8, enabling integer-only inference on supported hardware.
- **Runtime precision control** via `Option` flags (`use_fp16_storage`, `use_fp16_arithmetic`, `use_fp16_packed`) allows automatic type conversion during inference without model modification, with layers in [`src/layer/convolution.cpp`](https://github.com/Tencent/ncnn/blob/main/src/layer/convolution.cpp) and [`src/layer/innerproduct.cpp`](https://github.com/Tencent/ncnn/blob/main/src/layer/innerproduct.cpp) implementing the conversion logic.

## Frequently Asked Questions

### What data types does ncnn support for model inference?

ncnn supports **float32** (32-bit floating point), **float16** (16-bit half-precision), **bfloat16** (16-bit brain floating point), and **int8** (8-bit quantized integer). These types are stored in the `ncnn::Mat` container, which tracks the element size via `elemsize` (4 bytes for float32, 2 bytes for float16/bfloat16, and 1 byte for int8).

### How does ncnn convert between float32 and float16 during inference?

ncnn provides explicit conversion functions declared in [`src/mat.h`](https://github.com/Tencent/ncnn/blob/main/src/mat.h): `cast_float32_to_float16` and `cast_float16_to_float32`. These functions perform element-wise conversion using SIMD intrinsics optimized for the target architecture (NEON on ARM, AVX-512 on x86). During automatic inference, layers check `Option::use_fp16_storage` and `Option::use_fp16_arithmetic` to determine whether to cast weights and activations to float16 before computation.

### What is the difference between float16 and bfloat16 in ncnn?

**float16** (half-precision) uses 16 bits with a 5-bit exponent and 10-bit mantissa, providing higher precision but narrower range, ideal for memory-constrained mobile GPUs. **bfloat16** (brain floating point) uses 16 bits with an 8-bit exponent and 7-bit mantissa, matching float32's range but with reduced precision, optimized for AVX-512 CPUs. ncnn supports both via `cast_float32_to_float16`/`cast_float16_to_float32` and `float32_to_bfloat16`/`bfloat16_to_float32` respectively.

### How does ncnn handle int8 quantization for model weights and activations?

ncnn implements int8 quantization through the `quantize_to_int8` function in [`src/mat.cpp`](https://github.com/Tencent/ncnn/blob/main/src/mat.cpp) (line 1719), which converts float32 tensors to int8 using per-channel scale factors. During inference, layers like `InnerProduct` (line 73-74) and `Convolution` quantize weights during loading if int8 mode is enabled, while activations are quantized on-the-fly using pre-calculated scales stored in `scale_data` tensors. The `cast_int8_to_float32` function handles dequantization when layers require floating-point input.