How ncnn Int8 Quantization Works: Understanding use_int8_inference, use_int8_packed, use_int8_storage, and use_int8_arithmetic

ncnn's int8 quantization system is governed by four boolean flags in the Option class—use_int8_inference enables the int8 code path, use_int8_packed controls GPU tensor packing, use_int8_storage manages CPU blob format, and use_int8_arithmetic enables native int8 computation only when both packing and storage are enabled.

ncnn int8 quantization implements a post-training pipeline that converts floating-point models to 8-bit integer representations for efficient inference on mobile and embedded devices. The framework stores quantized weights and per-channel scales in the model file, then uses the four control flags declared in src/option.h to determine how activations are stored, packed, and computed at runtime.

Understanding the Four Int8 Control Flags

The behavior of ncnn int8 quantization is controlled by four boolean members of the Option structure defined in src/option.h at lines 81, 94, 95, and 96.

use_int8_inference: The Master Switch

The use_int8_inference flag acts as the global gate for the int8 code path. When set to true—the default value set in src/option.cpp at line 32—layers check whether their weights have been quantized to 8-bit before entering the optimized int8 branch.

In src/layer/convolution.cpp at line 189, the convolution layer verifies this flag before using int8 weights:

if (opt.use_int8_inference && weight_data.elemsize == (size_t)1u) { … }

If use_int8_inference is false, all layers execute their FP32 implementations regardless of weight quantization status.

use_int8_packed: GPU Memory Layout

The use_int8_packed flag controls whether int8 tensors use a packed format on the GPU. This flag defaults to true in src/option.cpp at line 40 and is primarily relevant for the Vulkan backend.

In src/gpu.cpp at lines 3416-3522, the GPU pipeline creation logic sets this flag based on hardware capabilities:

opt.use_int8_packed = use_int8;

The flag only takes effect if vkdev->info.support_int8_packed() returns true. When enabled, int8 tensors are stored in a layout optimized for GPU memory access patterns.

use_int8_storage: CPU Blob Format

The use_int8_storage flag determines whether intermediate activation blobs are stored as 8-bit integers or 32-bit floats on the CPU. Defaulting to true in src/option.cpp at line 41, this flag affects memory footprint and data movement costs.

When true, Mat objects use elemsize==1 for activations. When false, blobs remain in FP32 format even if weights are quantized, forcing dequantization at the layer boundary.

use_int8_arithmetic: Native 8-bit Computation

The use_int8_arithmetic flag enables actual int8 SIMD computations on the CPU, but it defaults to false in src/option.cpp at line 42 as a safety measure. This flag is automatically managed by the framework based on the other settings.

In src/net.cpp at lines 1072-1088, the network initialization enforces a strict dependency:

if (!opt.use_int8_packed && !opt.use_int8_storage) {
    opt.use_int8_arithmetic = false;
}

Int8 arithmetic is only enabled when both use_int8_packed and use_int8_storage are true. If either packing or storage is disabled, the system falls back to dequantize → FP32 compute → requantize, ensuring numerical stability at the cost of performance.

How the Flags Control Execution Flow

Layer-Level Gating

Every layer capable of quantization checks use_int8_inference before entering the optimized branch. As shown in src/layer/convolution.cpp at line 189, the layer verifies both the flag and that weights have been loaded with elemsize == 1:

if (opt.use_int8_inference && weight_data.elemsize == (size_t)1u) {
    // Execute INT8 optimized path
}

GPU Pipeline Validation

For Vulkan backends, src/gpu.cpp at lines 3416-3522 validates hardware support before enabling packed formats:

opt.use_int8_packed = use_int8;
opt.use_int8_storage = use_int8 && vkdev->info.support_int8_storage();

If vkdev->info.support_int8_packed() returns false, the framework clears use_int8_packed, which subsequently forces use_int8_arithmetic to false in src/net.cpp.

Arithmetic Fallback Paths

When use_int8_arithmetic is disabled, layers fall back to FP32 computation. In src/layer/x86/innerproduct_x86.cpp, the fast int8 path occupies lines 55-94, while the fallback dequantization path appears at lines 93-112:

// Fast path (lines 55-94)
if (opt.use_int8_arithmetic) {
    // Native INT8 SIMD operations
}

// Fallback path (lines 93-112)
// Dequantize to FP32, compute, requantize

The Post-Training Quantization Process

ncnn implements post-training quantization through the ncnn2int8 tool and runtime quantization layers.

Calibration and Scale Extraction

The tools/quantize/ncnn2int8.cpp utility performs calibration by running representative data through the model, collecting per-channel min/max statistics, and computing scale tensors (weight_data_int8_scales).

Runtime Weight Quantization

The quantize_to_int8 function in src/mat.cpp (lines 1919-1939) creates a Quantize layer to convert FP32 weights:

Layer* quantize = create_layer(LayerType::Quantize); // L1921
ParamDict pd;
pd.set(0, scale_data.w);                             // L1924-L1925
quantize->load_param(pd);
quantize->load_model(ModelBinFromMatArray(weights)); // L1929-L1931

Execution Strategy Selection

During inference, the network evaluates the four flags to select between:

  • Direct int8 arithmetic: Fastest path using SIMD instructions when use_int8_arithmetic is true
  • Dequantize-FP32-requantize: Fallback path when arithmetic is disabled but storage is int8
  • Full FP32: When use_int8_inference is false

Practical Configuration Examples

Disabling Int8 for Debugging

To verify model accuracy against FP32 baselines, disable the int8 inference path entirely:

#include "net.h"

ncnn::Net net;
net.opt.use_int8_inference = false;  // Force FP32 execution

net.load_param("model.param");
net.load_model("model.bin");

CPU Inference with Fallback

On CPUs lacking fast int8 SIMD, keep int8 weights for memory efficiency but force FP32 computation:

ncnn::Net net;
net.opt.use_int8_inference = true;
net.opt.use_int8_storage = true;      // Keep int8 blobs
net.opt.use_int8_arithmetic = false;  // Force dequantize -> FP32 -> requantize

GPU Inference with Vulkan

For Vulkan-capable devices, enable all int8 features for maximum throughput:

ncnn::Net net;
net.opt.use_int8_inference = true;
net.opt.use_int8_packed = true;    // Packed GPU format
net.opt.use_int8_storage = true;
// Arithmetic enabled automatically if hardware supports it

Python API Usage

The pybind11 bindings in python/src/main.cpp (lines 201-203) expose the same options:

import ncnn

net = ncnn.Net()
net.opt.use_int8_inference = True
net.opt.use_int8_storage = True
net.opt.use_int8_arithmetic = False  # Debug fallback

net.load_param("model_int8.param")
net.load_model("model_int8.bin")

Offline Model Quantization

Convert FP32 models to int8 using the calibration tool:

./ncnn2int8 \
    -i mobilenet_v2.param \
    -w mobilenet_v2.bin \
    -o mobilenet_v2_int8.param \
    -b mobilenet_v2_int8.bin \
    -p 8

Summary

  • use_int8_inference gates the entire int8 code path; when disabled, all layers execute FP32 kernels regardless of weight quantization.
  • use_int8_packed controls GPU-side tensor packing for Vulkan backends and requires hardware support via vkdev->info.support_int8_packed().
  • use_int8_storage determines whether CPU intermediate blobs remain as 8-bit integers or are stored as 32-bit floats.
  • use_int8_arithmetic enables native int8 SIMD computations but is automatically forced to false unless both packing and storage are enabled, as enforced in src/net.cpp lines 1072-1088.

Frequently Asked Questions

What happens if I set use_int8_inference to false?

Setting use_int8_inference to false forces ncnn to ignore all quantized weights and execute every layer using full FP32 arithmetic. As implemented in src/layer/convolution.cpp at line 189, layers check this flag before entering the int8 branch, meaning the FP32 code path runs even when 8-bit weights are present in memory. This configuration is primarily used for debugging accuracy issues or establishing FP32 baselines for comparison.

Why is use_int8_arithmetic false by default?

The use_int8_arithmetic flag defaults to false in src/option.cpp at line 42 to ensure numerical stability across diverse hardware platforms. According to the validation logic in src/net.cpp at lines 1072-1088, this flag is automatically cleared unless both use_int8_packed and use_int8_storage are enabled. This conservative default prevents performance degradation on CPUs lacking efficient int8 SIMD instructions, where the framework instead uses the dequantize → FP32 compute → requantize fallback path demonstrated in src/layer/x86/innerproduct_x86.cpp at lines 93-112.

How do I know if my GPU supports int8_packed?

GPU support for packed int8 formats is determined at runtime by querying Vulkan device capabilities in src/gpu.cpp at lines 3416-3522. The framework checks vkdev->info.support_int8_packed() and vkdev->info.support_int8_storage() during pipeline creation. If the device reports false for packed support, the framework clears use_int8_packed, which subsequently forces use_int8_arithmetic to false according to the logic in src/net.cpp. You can verify support by inspecting these capability bits in your Vulkan initialization logs or by checking whether the flags remain enabled after calling net.load_model().

Can I mix int8 weights with fp32 computation?

Yes, ncnn explicitly supports mixing int8 weights with FP32 computation through strategic flag configuration. By setting use_int8_inference = true to load quantized weights while setting use_int8_arithmetic = false, you force layers to execute the fallback path shown in src/layer/x86/innerproduct_x86.cpp at lines 93-112. This path dequantizes inputs to FP32, performs standard floating-point computation, and requantizes outputs if necessary. This configuration is particularly useful on CPUs without efficient int8 SIMD support, allowing you to benefit from reduced model size while maintaining computational accuracy through FP32 arithmetic.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →