# Understanding ncnn Mat Storage Formats: elemsize and elempack Explained

> Learn ncnn Mat storage formats elemsize and elempack optimize performance by defining element size and SIMD packing. Discover the best choices for your hardware and tensor dimensions.

- Repository: [Tencent/ncnn](https://github.com/tencent/ncnn)
- Tags: deep-dive
- Published: 2026-02-23

---

**The `ncnn::Mat` container uses `elemsize` to define the byte size of each element (precision) and `elempack` to specify how many elements are packed together for SIMD vectorization, with the optimal combination determined by hardware capabilities and tensor dimensions.**

The `Mat` class is the fundamental tensor data structure in Tencent/ncnn, a high-performance neural network inference framework optimized for mobile and edge devices. Understanding the `ncnn` Mat storage formats is essential for optimizing memory layout and computational throughput across CPU and Vulkan GPU backends. These two fields—defined in [`src/mat.h`](https://github.com/Tencent/ncnn/blob/main/src/mat.h)—determine whether data is stored as scalars or packed vectors, directly impacting inference speed.

## Core Storage Parameters

### elemsize: Element Precision

The `elemsize` field specifies the size in bytes of a single logical element. This value determines the numerical precision of the tensor:

- **4 bytes**: Standard FP32 (float) or INT32—default for most layers
- **2 bytes**: FP16 (half-precision)—optimal for NEON/AVX-FP16 acceleration  
- **1 byte**: INT8/UINT8—used for quantized models requiring minimal memory bandwidth

In [`src/mat.h`](https://github.com/Tencent/ncnn/blob/main/src/mat.h), `elemsize` is stored as a `size_t` member variable alongside the data pointer. Most constructors default to `4u` (4 bytes), but overloads accept explicit `elemsize` parameters for mixed-precision workflows.

### elempack: SIMD Vector Packing

The `elempack` field defines how many consecutive elements are stored contiguously to fill a SIMD register:

- **1**: Scalar format (no packing)—universal compatibility
- **4**: 128-bit SSE/NEON packs (four 32-bit floats or eight 16-bit values)
- **8**: 256-bit AVX/AVX2 packs—optimal for modern x86-64 processors
- **16**: 512-bit AVX-512 packs—maximum throughput on supported hardware

When `elempack > 1`, the framework interleaves data from consecutive channels so that a single vector instruction processes multiple elements simultaneously. The physical memory remains linear, but the logical stride (`cstep`) accounts for the packed layout.

## How NCNN Selects Optimal Packing

NCNN automatically determines the best `ncnn` Mat storage format during model loading and layer execution through three criteria:

1. **Hardware Capabilities**: At runtime, NCNN queries CPU features using functions like `cpu_support_sse2()`, `cpu_support_avx()`, `cpu_support_avx512f()`, and `cpu_support_neon()` to identify the widest SIMD registers available.

2. **Channel Divisibility**: The framework only applies packing when the channel dimension is evenly divisible by the pack size. If `c % 8 != 0` on AVX2 hardware, NCNN falls back to `elempack=4` or `elempack=1` to avoid partial vector processing.

3. **Layer Requirements**: Specific operations in `src/layer/` subdirectories call `convert_packing()` to transform inputs into their preferred layout. For example, depthwise convolutions may require specific packing to maximize vector utilization, prompting automatic conversion via the implementation in [`src/mat.cpp`](https://github.com/Tencent/ncnn/blob/main/src/mat.cpp).

## When to Use Each Storage Format

Selecting the appropriate `elemsize` and `elempack` combination depends on your model precision and target hardware:

**FP32 Models on x86-64 with AVX2**
- Use `elemsize=4` with `elempack=8` (when channels are multiples of 8)
- Falls back to `elempack=4` for SSE compatibility or `elempack=1` for odd channel counts
- AVX2 processes eight 32-bit floats per register, doubling throughput over scalar code

**FP16 Models on ARM NEON**
- Use `elemsize=2` with `elempack=4` (NEON 128-bit)
- NEON handles four half-precision values per instruction, offering 2x memory bandwidth reduction and significant speedup on mobile GPUs

**INT8 Quantized Inference**
- Use `elemsize=1` with `elempack=4` (SSE/NEON) or `elempack=8` (AVX2)
- Integer SIMD provides the highest performance gains for quantized models, where memory bandwidth is the primary bottleneck

**Vulkan GPU Backend**
- Use `elemsize=4` (or `2`/`1` based on precision requirements)
- The Vulkan backend mirrors CPU packing strategies, with `elempack` values handled by shader variants rather than CPU vector registers

## Working with Mat Storage Formats in Code

### Creating Packed Mats Manually

When you know the channel dimensions align with hardware capabilities, construct packed Mats directly using the overloaded constructor in [`src/mat.h`](https://github.com/Tencent/ncnn/blob/main/src/mat.h):

```cpp
// 32-bit float, packed 8-way for AVX2 (requires c % 8 == 0)
int w = 224, h = 224, c = 256;
ncnn::Mat packed_tensor(w, h, c, /*elemsize=*/4u, /*elempack=*/8);

```

For INT8 quantized tensors:

```cpp
// 8-bit integer, packed 4-way for NEON/SSE
ncnn::Mat int8_tensor(w, h, c, /*elemsize=*/1, /*elempack=*/4);

```

### Converting Between Packing Formats

Use the `convert_packing` function declared in [`src/mat.h`](https://github.com/Tencent/ncnn/blob/main/src/mat.h) and implemented in [`src/mat.cpp`](https://github.com/Tencent/ncnn/blob/main/src/mat.cpp) to transform existing Mats:

```cpp
ncnn::Mat src = ncnn::Mat::from_pixels(image_data, ncnn::Mat::PIXEL_RGB, w, h);
ncnn::Mat dst;
ncnn::Option opt;  // Uses current CPU features by default

// Convert to 4-way packing for SSE/NEON optimization
ncnn::convert_packing(src, dst, 4, opt);

```

This function handles the reordering of data from scalar to packed layout (or between different pack sizes) automatically based on the target `elempack` value.

### Inspecting Storage Properties

Verify the actual storage format of any Mat instance:

```cpp
printf("Storage: elemsize=%zu bytes, elempack=%d (pack ratio %dx)\n",
       mat.elemsize, mat.elempack, mat.elempack);
printf("Total bytes: %zu (cstep=%zu)\n", 
       mat.total() * mat.elemsize, mat.cstep);

```

## Summary

- **`elemsize`** controls precision (1, 2, or 4 bytes) and is defined in [`src/mat.h`](https://github.com/Tencent/ncnn/blob/main/src/mat.h) alongside the `Mat` class structure
- **`elempack`** enables SIMD vectorization (1, 4, 8, or 16) with automatic selection based on CPU capabilities and channel divisibility
- **Scalar format** (`elempack=1`) guarantees compatibility across all hardware but sacrifices vectorization performance
- **Packed formats** (`elempack>1`) maximize throughput when channel counts align with register widths (4 for 128-bit, 8 for 256-bit, 16 for 512-bit)
- **Conversion utilities** in [`src/mat.cpp`](https://github.com/Tencent/ncnn/blob/main/src/mat.cpp) provide `convert_packing()` for runtime layout transformations without manual data rearrangement

## Frequently Asked Questions

### What happens if I specify an elempack that doesn't divide the channel count evenly?

NCNN will still create the Mat, but operations may fall back to scalar processing or trigger assertions in debug builds. The framework automatically selects valid packing during layer execution via `convert_packing()`, so manual Mat creation should verify that `channels % elempack == 0` for optimal performance.

### Can I mix different elempack values within the same network?

Yes. Individual layers in `src/layer/` implementations frequently convert inputs to their preferred packing format before computation and convert back afterward. The `convert_packing()` function handles these transitions efficiently, allowing different layers to use different optimal pack sizes based on their specific SIMD requirements.

### Does the Vulkan backend use the same elempack values as the CPU?

Yes, the Vulkan backend uses identical `elempack` semantics (1, 4, 8) but implements packing through shader specialization constants rather than CPU vector registers. The `VkMat` class in [`src/mat.h`](https://github.com/Tencent/ncnn/blob/main/src/mat.h) mirrors the CPU `Mat` interface, using the same `elemsize` and `elempack` fields to describe GPU buffer layouts.

### How do I force scalar storage for debugging purposes?

Construct Mats with `elempack=1` or call `convert_packing(src, dst, 1, opt)` to unpack data. This removes all SIMD interleaving, presenting elements in standard row-major order without vector packing, which simplifies debugging and interoperability with external libraries.