# How ncnn Int8 Quantization Works: Understanding use_int8_inference, use_int8_packed, use_int8_storage, and use_int8_arithmetic

> Understand ncnn's int8 quantization with use_int8_inference, use_int8_packed, use_int8_storage, and use_int8_arithmetic flags. Optimize your inference performance.

- Repository: [Tencent/ncnn](https://github.com/tencent/ncnn)
- Tags: deep-dive
- Published: 2026-02-23

---

**ncnn's int8 quantization system is governed by four boolean flags in the `Option` class—`use_int8_inference` enables the int8 code path, `use_int8_packed` controls GPU tensor packing, `use_int8_storage` manages CPU blob format, and `use_int8_arithmetic` enables native int8 computation only when both packing and storage are enabled.**

ncnn int8 quantization implements a post-training pipeline that converts floating-point models to 8-bit integer representations for efficient inference on mobile and embedded devices. The framework stores quantized weights and per-channel scales in the model file, then uses the four control flags declared in [`src/option.h`](https://github.com/Tencent/ncnn/blob/main/src/option.h) to determine how activations are stored, packed, and computed at runtime.

## Understanding the Four Int8 Control Flags

The behavior of ncnn int8 quantization is controlled by four boolean members of the `Option` structure defined in [`src/option.h`](https://github.com/Tencent/ncnn/blob/main/src/option.h) at lines 81, 94, 95, and 96.

### use_int8_inference: The Master Switch

The `use_int8_inference` flag acts as the global gate for the int8 code path. When set to `true`—the default value set in [`src/option.cpp`](https://github.com/Tencent/ncnn/blob/main/src/option.cpp) at line 32—layers check whether their weights have been quantized to 8-bit before entering the optimized int8 branch.

In [`src/layer/convolution.cpp`](https://github.com/Tencent/ncnn/blob/main/src/layer/convolution.cpp) at line 189, the convolution layer verifies this flag before using int8 weights:

```cpp
if (opt.use_int8_inference && weight_data.elemsize == (size_t)1u) { … }

```

If `use_int8_inference` is `false`, all layers execute their FP32 implementations regardless of weight quantization status.

### use_int8_packed: GPU Memory Layout

The `use_int8_packed` flag controls whether int8 tensors use a packed format on the GPU. This flag defaults to `true` in [`src/option.cpp`](https://github.com/Tencent/ncnn/blob/main/src/option.cpp) at line 40 and is primarily relevant for the Vulkan backend.

In [`src/gpu.cpp`](https://github.com/Tencent/ncnn/blob/main/src/gpu.cpp) at lines 3416-3522, the GPU pipeline creation logic sets this flag based on hardware capabilities:

```cpp
opt.use_int8_packed = use_int8;

```

The flag only takes effect if `vkdev->info.support_int8_packed()` returns true. When enabled, int8 tensors are stored in a layout optimized for GPU memory access patterns.

### use_int8_storage: CPU Blob Format

The `use_int8_storage` flag determines whether intermediate activation blobs are stored as 8-bit integers or 32-bit floats on the CPU. Defaulting to `true` in [`src/option.cpp`](https://github.com/Tencent/ncnn/blob/main/src/option.cpp) at line 41, this flag affects memory footprint and data movement costs.

When `true`, `Mat` objects use `elemsize==1` for activations. When `false`, blobs remain in FP32 format even if weights are quantized, forcing dequantization at the layer boundary.

### use_int8_arithmetic: Native 8-bit Computation

The `use_int8_arithmetic` flag enables actual int8 SIMD computations on the CPU, but it defaults to `false` in [`src/option.cpp`](https://github.com/Tencent/ncnn/blob/main/src/option.cpp) at line 42 as a safety measure. This flag is automatically managed by the framework based on the other settings.

In [`src/net.cpp`](https://github.com/Tencent/ncnn/blob/main/src/net.cpp) at lines 1072-1088, the network initialization enforces a strict dependency:

```cpp
if (!opt.use_int8_packed && !opt.use_int8_storage) {
    opt.use_int8_arithmetic = false;
}

```

**Int8 arithmetic is only enabled when both `use_int8_packed` and `use_int8_storage` are true.** If either packing or storage is disabled, the system falls back to dequantize → FP32 compute → requantize, ensuring numerical stability at the cost of performance.

## How the Flags Control Execution Flow

### Layer-Level Gating

Every layer capable of quantization checks `use_int8_inference` before entering the optimized branch. As shown in [`src/layer/convolution.cpp`](https://github.com/Tencent/ncnn/blob/main/src/layer/convolution.cpp) at line 189, the layer verifies both the flag and that weights have been loaded with `elemsize == 1`:

```cpp
if (opt.use_int8_inference && weight_data.elemsize == (size_t)1u) {
    // Execute INT8 optimized path
}

```

### GPU Pipeline Validation

For Vulkan backends, [`src/gpu.cpp`](https://github.com/Tencent/ncnn/blob/main/src/gpu.cpp) at lines 3416-3522 validates hardware support before enabling packed formats:

```cpp
opt.use_int8_packed = use_int8;
opt.use_int8_storage = use_int8 && vkdev->info.support_int8_storage();

```

If `vkdev->info.support_int8_packed()` returns false, the framework clears `use_int8_packed`, which subsequently forces `use_int8_arithmetic` to false in [`src/net.cpp`](https://github.com/Tencent/ncnn/blob/main/src/net.cpp).

### Arithmetic Fallback Paths

When `use_int8_arithmetic` is disabled, layers fall back to FP32 computation. In [`src/layer/x86/innerproduct_x86.cpp`](https://github.com/Tencent/ncnn/blob/main/src/layer/x86/innerproduct_x86.cpp), the fast int8 path occupies lines 55-94, while the fallback dequantization path appears at lines 93-112:

```cpp
// Fast path (lines 55-94)
if (opt.use_int8_arithmetic) {
    // Native INT8 SIMD operations
}

// Fallback path (lines 93-112)
// Dequantize to FP32, compute, requantize

```

## The Post-Training Quantization Process

ncnn implements post-training quantization through the `ncnn2int8` tool and runtime quantization layers.

### Calibration and Scale Extraction

The [`tools/quantize/ncnn2int8.cpp`](https://github.com/Tencent/ncnn/blob/main/tools/quantize/ncnn2int8.cpp) utility performs calibration by running representative data through the model, collecting per-channel min/max statistics, and computing scale tensors (`weight_data_int8_scales`).

### Runtime Weight Quantization

The `quantize_to_int8` function in [`src/mat.cpp`](https://github.com/Tencent/ncnn/blob/main/src/mat.cpp) (lines 1919-1939) creates a `Quantize` layer to convert FP32 weights:

```cpp
Layer* quantize = create_layer(LayerType::Quantize); // L1921
ParamDict pd;
pd.set(0, scale_data.w);                             // L1924-L1925
quantize->load_param(pd);
quantize->load_model(ModelBinFromMatArray(weights)); // L1929-L1931

```

### Execution Strategy Selection

During inference, the network evaluates the four flags to select between:
- **Direct int8 arithmetic**: Fastest path using SIMD instructions when `use_int8_arithmetic` is true
- **Dequantize-FP32-requantize**: Fallback path when arithmetic is disabled but storage is int8
- **Full FP32**: When `use_int8_inference` is false

## Practical Configuration Examples

### Disabling Int8 for Debugging

To verify model accuracy against FP32 baselines, disable the int8 inference path entirely:

```cpp
#include "net.h"

ncnn::Net net;
net.opt.use_int8_inference = false;  // Force FP32 execution

net.load_param("model.param");
net.load_model("model.bin");

```

### CPU Inference with Fallback

On CPUs lacking fast int8 SIMD, keep int8 weights for memory efficiency but force FP32 computation:

```cpp
ncnn::Net net;
net.opt.use_int8_inference = true;
net.opt.use_int8_storage = true;      // Keep int8 blobs
net.opt.use_int8_arithmetic = false;  // Force dequantize -> FP32 -> requantize

```

### GPU Inference with Vulkan

For Vulkan-capable devices, enable all int8 features for maximum throughput:

```cpp
ncnn::Net net;
net.opt.use_int8_inference = true;
net.opt.use_int8_packed = true;    // Packed GPU format
net.opt.use_int8_storage = true;
// Arithmetic enabled automatically if hardware supports it

```

### Python API Usage

The pybind11 bindings in [`python/src/main.cpp`](https://github.com/Tencent/ncnn/blob/main/python/src/main.cpp) (lines 201-203) expose the same options:

```python
import ncnn

net = ncnn.Net()
net.opt.use_int8_inference = True
net.opt.use_int8_storage = True
net.opt.use_int8_arithmetic = False  # Debug fallback

net.load_param("model_int8.param")
net.load_model("model_int8.bin")

```

### Offline Model Quantization

Convert FP32 models to int8 using the calibration tool:

```bash
./ncnn2int8 \
    -i mobilenet_v2.param \
    -w mobilenet_v2.bin \
    -o mobilenet_v2_int8.param \
    -b mobilenet_v2_int8.bin \
    -p 8

```

## Summary

- **`use_int8_inference`** gates the entire int8 code path; when disabled, all layers execute FP32 kernels regardless of weight quantization.
- **`use_int8_packed`** controls GPU-side tensor packing for Vulkan backends and requires hardware support via `vkdev->info.support_int8_packed()`.
- **`use_int8_storage`** determines whether CPU intermediate blobs remain as 8-bit integers or are stored as 32-bit floats.
- **`use_int8_arithmetic`** enables native int8 SIMD computations but is automatically forced to `false` unless both packing and storage are enabled, as enforced in [`src/net.cpp`](https://github.com/Tencent/ncnn/blob/main/src/net.cpp) lines 1072-1088.

## Frequently Asked Questions

### What happens if I set use_int8_inference to false?

Setting `use_int8_inference` to `false` forces ncnn to ignore all quantized weights and execute every layer using full FP32 arithmetic. As implemented in [`src/layer/convolution.cpp`](https://github.com/Tencent/ncnn/blob/main/src/layer/convolution.cpp) at line 189, layers check this flag before entering the int8 branch, meaning the FP32 code path runs even when 8-bit weights are present in memory. This configuration is primarily used for debugging accuracy issues or establishing FP32 baselines for comparison.

### Why is use_int8_arithmetic false by default?

The `use_int8_arithmetic` flag defaults to `false` in [`src/option.cpp`](https://github.com/Tencent/ncnn/blob/main/src/option.cpp) at line 42 to ensure numerical stability across diverse hardware platforms. According to the validation logic in [`src/net.cpp`](https://github.com/Tencent/ncnn/blob/main/src/net.cpp) at lines 1072-1088, this flag is automatically cleared unless both `use_int8_packed` and `use_int8_storage` are enabled. This conservative default prevents performance degradation on CPUs lacking efficient int8 SIMD instructions, where the framework instead uses the dequantize → FP32 compute → requantize fallback path demonstrated in [`src/layer/x86/innerproduct_x86.cpp`](https://github.com/Tencent/ncnn/blob/main/src/layer/x86/innerproduct_x86.cpp) at lines 93-112.

### How do I know if my GPU supports int8_packed?

GPU support for packed int8 formats is determined at runtime by querying Vulkan device capabilities in [`src/gpu.cpp`](https://github.com/Tencent/ncnn/blob/main/src/gpu.cpp) at lines 3416-3522. The framework checks `vkdev->info.support_int8_packed()` and `vkdev->info.support_int8_storage()` during pipeline creation. If the device reports false for packed support, the framework clears `use_int8_packed`, which subsequently forces `use_int8_arithmetic` to false according to the logic in [`src/net.cpp`](https://github.com/Tencent/ncnn/blob/main/src/net.cpp). You can verify support by inspecting these capability bits in your Vulkan initialization logs or by checking whether the flags remain enabled after calling `net.load_model()`.

### Can I mix int8 weights with fp32 computation?

Yes, ncnn explicitly supports mixing int8 weights with FP32 computation through strategic flag configuration. By setting `use_int8_inference = true` to load quantized weights while setting `use_int8_arithmetic = false`, you force layers to execute the fallback path shown in [`src/layer/x86/innerproduct_x86.cpp`](https://github.com/Tencent/ncnn/blob/main/src/layer/x86/innerproduct_x86.cpp) at lines 93-112. This path dequantizes inputs to FP32, performs standard floating-point computation, and requantizes outputs if necessary. This configuration is particularly useful on CPUs without efficient int8 SIMD support, allowing you to benefit from reduced model size while maintaining computational accuracy through FP32 arithmetic.