Understanding ncnn Mat Storage Formats: elemsize and elempack Explained

The ncnn::Mat container uses elemsize to define the byte size of each element (precision) and elempack to specify how many elements are packed together for SIMD vectorization, with the optimal combination determined by hardware capabilities and tensor dimensions.

The Mat class is the fundamental tensor data structure in Tencent/ncnn, a high-performance neural network inference framework optimized for mobile and edge devices. Understanding the ncnn Mat storage formats is essential for optimizing memory layout and computational throughput across CPU and Vulkan GPU backends. These two fields—defined in src/mat.h—determine whether data is stored as scalars or packed vectors, directly impacting inference speed.

Core Storage Parameters

elemsize: Element Precision

The elemsize field specifies the size in bytes of a single logical element. This value determines the numerical precision of the tensor:

  • 4 bytes: Standard FP32 (float) or INT32—default for most layers
  • 2 bytes: FP16 (half-precision)—optimal for NEON/AVX-FP16 acceleration
  • 1 byte: INT8/UINT8—used for quantized models requiring minimal memory bandwidth

In src/mat.h, elemsize is stored as a size_t member variable alongside the data pointer. Most constructors default to 4u (4 bytes), but overloads accept explicit elemsize parameters for mixed-precision workflows.

elempack: SIMD Vector Packing

The elempack field defines how many consecutive elements are stored contiguously to fill a SIMD register:

  • 1: Scalar format (no packing)—universal compatibility
  • 4: 128-bit SSE/NEON packs (four 32-bit floats or eight 16-bit values)
  • 8: 256-bit AVX/AVX2 packs—optimal for modern x86-64 processors
  • 16: 512-bit AVX-512 packs—maximum throughput on supported hardware

When elempack > 1, the framework interleaves data from consecutive channels so that a single vector instruction processes multiple elements simultaneously. The physical memory remains linear, but the logical stride (cstep) accounts for the packed layout.

How NCNN Selects Optimal Packing

NCNN automatically determines the best ncnn Mat storage format during model loading and layer execution through three criteria:

  1. Hardware Capabilities: At runtime, NCNN queries CPU features using functions like cpu_support_sse2(), cpu_support_avx(), cpu_support_avx512f(), and cpu_support_neon() to identify the widest SIMD registers available.

  2. Channel Divisibility: The framework only applies packing when the channel dimension is evenly divisible by the pack size. If c % 8 != 0 on AVX2 hardware, NCNN falls back to elempack=4 or elempack=1 to avoid partial vector processing.

  3. Layer Requirements: Specific operations in src/layer/ subdirectories call convert_packing() to transform inputs into their preferred layout. For example, depthwise convolutions may require specific packing to maximize vector utilization, prompting automatic conversion via the implementation in src/mat.cpp.

When to Use Each Storage Format

Selecting the appropriate elemsize and elempack combination depends on your model precision and target hardware:

FP32 Models on x86-64 with AVX2

  • Use elemsize=4 with elempack=8 (when channels are multiples of 8)
  • Falls back to elempack=4 for SSE compatibility or elempack=1 for odd channel counts
  • AVX2 processes eight 32-bit floats per register, doubling throughput over scalar code

FP16 Models on ARM NEON

  • Use elemsize=2 with elempack=4 (NEON 128-bit)
  • NEON handles four half-precision values per instruction, offering 2x memory bandwidth reduction and significant speedup on mobile GPUs

INT8 Quantized Inference

  • Use elemsize=1 with elempack=4 (SSE/NEON) or elempack=8 (AVX2)
  • Integer SIMD provides the highest performance gains for quantized models, where memory bandwidth is the primary bottleneck

Vulkan GPU Backend

  • Use elemsize=4 (or 2/1 based on precision requirements)
  • The Vulkan backend mirrors CPU packing strategies, with elempack values handled by shader variants rather than CPU vector registers

Working with Mat Storage Formats in Code

Creating Packed Mats Manually

When you know the channel dimensions align with hardware capabilities, construct packed Mats directly using the overloaded constructor in src/mat.h:

// 32-bit float, packed 8-way for AVX2 (requires c % 8 == 0)
int w = 224, h = 224, c = 256;
ncnn::Mat packed_tensor(w, h, c, /*elemsize=*/4u, /*elempack=*/8);

For INT8 quantized tensors:

// 8-bit integer, packed 4-way for NEON/SSE
ncnn::Mat int8_tensor(w, h, c, /*elemsize=*/1, /*elempack=*/4);

Converting Between Packing Formats

Use the convert_packing function declared in src/mat.h and implemented in src/mat.cpp to transform existing Mats:

ncnn::Mat src = ncnn::Mat::from_pixels(image_data, ncnn::Mat::PIXEL_RGB, w, h);
ncnn::Mat dst;
ncnn::Option opt;  // Uses current CPU features by default

// Convert to 4-way packing for SSE/NEON optimization
ncnn::convert_packing(src, dst, 4, opt);

This function handles the reordering of data from scalar to packed layout (or between different pack sizes) automatically based on the target elempack value.

Inspecting Storage Properties

Verify the actual storage format of any Mat instance:

printf("Storage: elemsize=%zu bytes, elempack=%d (pack ratio %dx)\n",
       mat.elemsize, mat.elempack, mat.elempack);
printf("Total bytes: %zu (cstep=%zu)\n", 
       mat.total() * mat.elemsize, mat.cstep);

Summary

  • elemsize controls precision (1, 2, or 4 bytes) and is defined in src/mat.h alongside the Mat class structure
  • elempack enables SIMD vectorization (1, 4, 8, or 16) with automatic selection based on CPU capabilities and channel divisibility
  • Scalar format (elempack=1) guarantees compatibility across all hardware but sacrifices vectorization performance
  • Packed formats (elempack>1) maximize throughput when channel counts align with register widths (4 for 128-bit, 8 for 256-bit, 16 for 512-bit)
  • Conversion utilities in src/mat.cpp provide convert_packing() for runtime layout transformations without manual data rearrangement

Frequently Asked Questions

What happens if I specify an elempack that doesn't divide the channel count evenly?

NCNN will still create the Mat, but operations may fall back to scalar processing or trigger assertions in debug builds. The framework automatically selects valid packing during layer execution via convert_packing(), so manual Mat creation should verify that channels % elempack == 0 for optimal performance.

Can I mix different elempack values within the same network?

Yes. Individual layers in src/layer/ implementations frequently convert inputs to their preferred packing format before computation and convert back afterward. The convert_packing() function handles these transitions efficiently, allowing different layers to use different optimal pack sizes based on their specific SIMD requirements.

Does the Vulkan backend use the same elempack values as the CPU?

Yes, the Vulkan backend uses identical elempack semantics (1, 4, 8) but implements packing through shader specialization constants rather than CPU vector registers. The VkMat class in src/mat.h mirrors the CPU Mat interface, using the same elemsize and elempack fields to describe GPU buffer layouts.

How do I force scalar storage for debugging purposes?

Construct Mats with elempack=1 or call convert_packing(src, dst, 1, opt) to unpack data. This removes all SIMD interleaving, presenting elements in standard row-major order without vector packing, which simplifies debugging and interoperability with external libraries.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →