Comparing ncnn Convolution Implementations: SGEMM vs Winograd vs Standard

Tencent's ncnn library provides three distinct convolution implementations—standard direct convolution, SGEMM-based matrix multiplication, and Winograd minimal filtering—each selected automatically or manually via Option flags to optimize inference speed across different kernel sizes and tensor dimensions.

The ncnn inference framework optimizes deep learning deployments on mobile and edge devices by offering multiple algorithmic paths for the convolution operation. Understanding the differences between standard convolution, convolution_sgemm, and convolution_winograd enables developers to tune the ncnn::Option configuration for maximum throughput on specific hardware.

Overview of ncnn Convolution Strategies

The three implementations represent different trade-offs between arithmetic complexity, memory bandwidth, and temporary buffer usage.

Implementation Algorithm Selection Criteria Optimal Scenario
Standard Direct nested loops over input, kernel, and output channels use_sgemm_convolution=false and use_winograd_convolution=false Small tensors, 1×1 kernels, memory-constrained environments
SGEMM Im2col transformation followed by matrix multiplication use_sgemm_convolution=true and prefer_sgemm heuristic returns true Large kernels, many channels, when CPU vector units (AVX/NEON) can be saturated
Winograd Winograd minimal filtering (F(2×2, 3) for 3×3 kernels) use_winograd_convolution=true, kernel=3×3, stride=1, dilation=1, and prefer_winograd holds 3×3 convolutions with ≥8 input and output channels

Standard Convolution Implementation

The baseline implementation in src/layer/convolution.cpp performs a direct multiply-accumulate for every output element without transformation overhead.

static int convolution(const Mat& bottom_blob, Mat& top_blob,
                      const Mat& weight_data, const Mat& bias_data,
                      int kernel_w, int kernel_h, int stride_w, int stride_h,
                      int dilation_w, int dilation_h, int activation_type,
                      const Mat& activation_params, const Option& opt)
{
    // Nested loops: output channels → height → width
    // Inner loops: input channels × kernel_h × kernel_w
    // Accumulation: sum += input_val * weight_val
}

This path minimizes memory usage by avoiding the temporary buffers required for im2col or Winograd transforms. However, it fails to exploit cache-friendly memory access patterns or CPU vector instructions, making it slower for large-scale convolutions.

SGEMM-Based Convolution

The convolution_sgemm path transforms the convolution into a general matrix multiplication (GEMM) problem, enabling highly optimized BLAS-like kernels.

Im2col Transformation

The implementation in src/layer/x86/convolution_im2col_gemm.h (and ARM equivalent) first packs the input feature map into a column-major matrix:

// Transform kernel for SGEMM
convolution_im2col_gemm_transform_kernel(...);

// Pack input via im2col
convolution_im2col_pack_A_tile(...);

Heuristic Selection

The dispatcher in src/layer/x86/convolution_x86.cpp uses cache-size heuristics to decide between standard and SGEMM paths:

bool prefer_sgemm = num_input * num_output * kernel_w * kernel_h *
                    dilation_w * dilation_h * stride_w * stride_h *
                    (int)sizeof(float) * 2 > l2_cache_size ||
                    (num_input > 16 || num_output > 16);

if ((opt.use_sgemm_convolution && prefer_sgemm) || (kernel_w == 1 && kernel_h == 1))
{
    // Execute SGEMM path
}

SGEMM excels when the arithmetic intensity is high enough to amortize the im2col overhead and fully utilize AVX2/AVX-512 or NEON vector units.

Winograd Convolution

For 3×3 kernels with unit stride and dilation, convolution_winograd implements the F(2×2, 3) algorithm, reducing the multiplication count by approximately 2× at the cost of transformation overhead.

Transformation Pipeline

Located in src/layer/x86/convolution_3x3_winograd.h, the implementation follows the Winograd minimal filtering pattern:

// 1. Transform kernel to Winograd domain
convolution_3x3_winograd_transform_kernel(...);

// 2. Transform input tiles (B-matrix multiplication)
convolution_3x3_winograd_transform_input(...);

// 3. Element-wise multiply in Winograd domain

// 4. Inverse transform to spatial output (A-matrix multiplication)
convolution_3x3_winograd_transform_output(...);

Activation Conditions

The dispatcher enables Winograd only under specific constraints:

bool prefer_winograd = num_input >= 8 && num_output >= 8;

if ((opt.use_winograd_convolution && prefer_winograd &&
     kernel_w == 3 && kernel_h == 3 && 
     dilation_w == 1 && dilation_h == 1 &&
     stride_w == 1 && stride_h == 1) || 
    (!opt.use_sgemm_convolution))
{
    // Winograd path selected
}

Winograd delivers optimal throughput for 3×3 convolutions in ResNet-style architectures where channel counts exceed 8, as the transform overhead becomes negligible compared to the O(n²) arithmetic reduction.

When to Use Each Implementation

Use this decision matrix to configure ncnn::Option for your deployment:

Scenario Recommended Setting Rationale
1×1 convolutions use_sgemm_convolution = true (default) Automatically selected; GEMM is optimal for pointwise ops.
3×3 stride-1, channels ≥ 8 use_winograd_convolution = true ~2× speedup via Winograd minimal filtering.
Large kernels (5×5, 7×7) or high resolution use_sgemm_convolution = true Im2col + GEMM saturates vector units better than direct loops.
Memory-constrained embedded devices use_sgemm_convolution = falseuse_winograd_convolution = false Avoids temporary buffers for im2col and Winograd transforms.
Quantized INT8 inference Enable NCNN_INT8 compile flag; use same heuristics Both SGEMM and Winograd have INT8 specializations in *_int8.h files.

Practical Code Examples

Enabling SGEMM Convolution

Force the matrix-multiplication path for large models:

#include "net.h"

int main()
{
    ncnn::Option opt;
    opt.use_sgemm_convolution = true;  // Enable im2col + GEMM
    opt.use_winograd_convolution = false; // Disable Winograd fallback
    opt.num_threads = 4;

    ncnn::Net net;
    net.opt = opt;
    net.load_param("resnet50.param");
    net.load_model("resnet50.bin");
    
    // Inference continues...
}

Enabling Winograd Convolution

Optimize 3×3 layers in ResNet-style architectures:

ncnn::Option opt;
opt.use_winograd_convolution = true;  // Enable F(2x2,3) algorithm
opt.use_sgemm_convolution = true;     // Allow SGEMM for non-3x3 layers

ncnn::Net net;
net.opt = opt;
// Load model...

Disabling Optimizations for Debugging

Force the naive implementation to verify numerical correctness:

ncnn::Option opt;
opt.use_sgemm_convolution = false;
opt.use_winograd_convolution = false;
// Now uses src/layer/convolution.cpp directly

Key Source Files

File Path Description
src/layer/convolution.cpp Baseline direct convolution with nested loops over input, kernel, and output channels.
src/layer/x86/convolution_x86.cpp x86 dispatcher that implements prefer_sgemm and prefer_winograd heuristics to select the optimal path.
src/layer/x86/convolution_im2col_gemm.h Im2col packing and kernel transformation for SGEMM-based convolution on x86 (AVX/AVX2/AVX-512).
src/layer/arm/convolution_im2col_gemm.h ARM NEON implementation of im2col SGEMM for mobile/embedded devices.
src/layer/x86/convolution_3x3_winograd.h Winograd F(2×2, 3) transforms and tile-level computation for 3×3 kernels on x86.
src/layer/arm/convolution_3x3_winograd.h ARM NEON Winograd implementation for efficient 3×3 convolution on mobile CPUs.
src/option.h Runtime configuration struct containing use_sgemm_convolution and use_winograd_convolution flags.

Summary

  • Standard convolution in src/layer/convolution.cpp provides a memory-efficient baseline using direct nested loops, ideal for small tensors or when optimization buffers are prohibitive.
  • SGEMM convolution transforms inputs via im2col and executes high-performance matrix multiplication, automatically selected when prefer_sgemm heuristics detect large channel counts or spatial dimensions that benefit from AVX/AVX-512 or NEON vectorization.
  • Winograd convolution reduces arithmetic operations by ~2× for 3×3 kernels using the F(2×2, 3) algorithm, activated when use_winograd_convolution is enabled and channel counts exceed 8, making it optimal for ResNet-style architectures.

Frequently Asked Questions

What is the default convolution implementation in ncnn?

By default, ncnn enables SGEMM-based convolution (opt.use_sgemm_convolution is true in src/option.cpp). The dispatcher in src/layer/x86/convolution_x86.cpp evaluates the prefer_sgemm heuristic based on L2 cache size and channel dimensions, falling back to standard convolution only when SGEMM is disabled or suboptimal.

When should I disable Winograd convolution?

Disable Winograd convolution by setting opt.use_winograd_convolution = false when running models with non-3×3 kernels, strides greater than 1, or dilation factors other than 1, as the Winograd implementation in src/layer/x86/convolution_3x3_winograd.h strictly requires 3×3 kernels with unit stride and dilation. Additionally, disable it for models with fewer than 8 input or output channels where transform overhead exceeds the arithmetic savings.

How does ncnn decide between SGEMM and Winograd?

The decision logic in src/layer/x86/convolution_x86.cpp evaluates boolean flags prefer_sgemm and prefer_winograd. Winograd is selected when use_winograd_convolution is true, the kernel is 3×3 with stride=1 and dilation=1, and prefer_winograd (channels ≥ 8) holds true. SGEMM is selected when use_sgemm_convolution is true and the tensor size exceeds L2 cache capacity or channels exceed 16, unless Winograd conditions are met and enabled.

Can I use these optimized convolutions with INT8 quantization?

Yes, both SGEMM and Winograd implementations provide INT8 specializations found in corresponding *_int8.h files (e.g., src/layer/x86/convolution_im2col_gemm_int8.h and src/layer/x86/convolution_3x3_winograd_int8.h). To utilize them, compile ncnn with the NCNN_INT8 macro enabled and apply the same heuristic-based selection via Option flags, ensuring that the quantized model weights are properly calibrated for the chosen algorithmic path.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →