# Comparing ncnn Convolution Implementations: SGEMM vs Winograd vs Standard

> Explore ncnn's convolution implementations: SGEMM, Winograd, and standard. Discover how to optimize inference speed by choosing the best convolution method for your tensor dimensions and kernel sizes.

- Repository: [Tencent/ncnn](https://github.com/tencent/ncnn)
- Tags: performance
- Published: 2026-02-23

---

**Tencent's ncnn library provides three distinct convolution implementations—standard direct convolution, SGEMM-based matrix multiplication, and Winograd minimal filtering—each selected automatically or manually via `Option` flags to optimize inference speed across different kernel sizes and tensor dimensions.**

The ncnn inference framework optimizes deep learning deployments on mobile and edge devices by offering multiple algorithmic paths for the convolution operation. Understanding the differences between **standard convolution**, **convolution_sgemm**, and **convolution_winograd** enables developers to tune the `ncnn::Option` configuration for maximum throughput on specific hardware.

## Overview of ncnn Convolution Strategies

The three implementations represent different trade-offs between arithmetic complexity, memory bandwidth, and temporary buffer usage.

| Implementation | Algorithm | Selection Criteria | Optimal Scenario |
|----------------|-----------|-------------------|------------------|
| **Standard** | Direct nested loops over input, kernel, and output channels | `use_sgemm_convolution=false` and `use_winograd_convolution=false` | Small tensors, 1×1 kernels, memory-constrained environments |
| **SGEMM** | Im2col transformation followed by matrix multiplication | `use_sgemm_convolution=true` and `prefer_sgemm` heuristic returns true | Large kernels, many channels, when CPU vector units (AVX/NEON) can be saturated |
| **Winograd** | Winograd minimal filtering (F(2×2, 3) for 3×3 kernels) | `use_winograd_convolution=true`, kernel=3×3, stride=1, dilation=1, and `prefer_winograd` holds | 3×3 convolutions with ≥8 input and output channels |

## Standard Convolution Implementation

The baseline implementation in [`src/layer/convolution.cpp`](https://github.com/Tencent/ncnn/blob/main/src/layer/convolution.cpp) performs a direct multiply-accumulate for every output element without transformation overhead.

```cpp
static int convolution(const Mat& bottom_blob, Mat& top_blob,
                      const Mat& weight_data, const Mat& bias_data,
                      int kernel_w, int kernel_h, int stride_w, int stride_h,
                      int dilation_w, int dilation_h, int activation_type,
                      const Mat& activation_params, const Option& opt)
{
    // Nested loops: output channels → height → width
    // Inner loops: input channels × kernel_h × kernel_w
    // Accumulation: sum += input_val * weight_val
}

```

This path minimizes memory usage by avoiding the temporary buffers required for im2col or Winograd transforms. However, it fails to exploit cache-friendly memory access patterns or CPU vector instructions, making it slower for large-scale convolutions.

## SGEMM-Based Convolution

The `convolution_sgemm` path transforms the convolution into a general matrix multiplication (GEMM) problem, enabling highly optimized BLAS-like kernels.

### Im2col Transformation

The implementation in [`src/layer/x86/convolution_im2col_gemm.h`](https://github.com/Tencent/ncnn/blob/main/src/layer/x86/convolution_im2col_gemm.h) (and ARM equivalent) first packs the input feature map into a column-major matrix:

```cpp
// Transform kernel for SGEMM
convolution_im2col_gemm_transform_kernel(...);

// Pack input via im2col
convolution_im2col_pack_A_tile(...);

```

### Heuristic Selection

The dispatcher in [`src/layer/x86/convolution_x86.cpp`](https://github.com/Tencent/ncnn/blob/main/src/layer/x86/convolution_x86.cpp) uses cache-size heuristics to decide between standard and SGEMM paths:

```cpp
bool prefer_sgemm = num_input * num_output * kernel_w * kernel_h *
                    dilation_w * dilation_h * stride_w * stride_h *
                    (int)sizeof(float) * 2 > l2_cache_size ||
                    (num_input > 16 || num_output > 16);

if ((opt.use_sgemm_convolution && prefer_sgemm) || (kernel_w == 1 && kernel_h == 1))
{
    // Execute SGEMM path
}

```

**SGEMM excels** when the arithmetic intensity is high enough to amortize the im2col overhead and fully utilize AVX2/AVX-512 or NEON vector units.

## Winograd Convolution

For 3×3 kernels with unit stride and dilation, `convolution_winograd` implements the F(2×2, 3) algorithm, reducing the multiplication count by approximately 2× at the cost of transformation overhead.

### Transformation Pipeline

Located in [`src/layer/x86/convolution_3x3_winograd.h`](https://github.com/Tencent/ncnn/blob/main/src/layer/x86/convolution_3x3_winograd.h), the implementation follows the Winograd minimal filtering pattern:

```cpp
// 1. Transform kernel to Winograd domain
convolution_3x3_winograd_transform_kernel(...);

// 2. Transform input tiles (B-matrix multiplication)
convolution_3x3_winograd_transform_input(...);

// 3. Element-wise multiply in Winograd domain

// 4. Inverse transform to spatial output (A-matrix multiplication)
convolution_3x3_winograd_transform_output(...);

```

### Activation Conditions

The dispatcher enables Winograd only under specific constraints:

```cpp
bool prefer_winograd = num_input >= 8 && num_output >= 8;

if ((opt.use_winograd_convolution && prefer_winograd &&
     kernel_w == 3 && kernel_h == 3 && 
     dilation_w == 1 && dilation_h == 1 &&
     stride_w == 1 && stride_h == 1) || 
    (!opt.use_sgemm_convolution))
{
    // Winograd path selected
}

```

**Winograd delivers** optimal throughput for 3×3 convolutions in ResNet-style architectures where channel counts exceed 8, as the transform overhead becomes negligible compared to the O(n²) arithmetic reduction.

## When to Use Each Implementation

Use this decision matrix to configure `ncnn::Option` for your deployment:

| Scenario | Recommended Setting | Rationale |
|----------|---------------------|-----------|
| **1×1 convolutions** | `use_sgemm_convolution = true` (default) | Automatically selected; GEMM is optimal for pointwise ops. |
| **3×3 stride-1, channels ≥ 8** | `use_winograd_convolution = true` | ~2× speedup via Winograd minimal filtering. |
| **Large kernels (5×5, 7×7) or high resolution** | `use_sgemm_convolution = true` | Im2col + GEMM saturates vector units better than direct loops. |
| **Memory-constrained embedded devices** | `use_sgemm_convolution = false`<br>`use_winograd_convolution = false` | Avoids temporary buffers for im2col and Winograd transforms. |
| **Quantized INT8 inference** | Enable `NCNN_INT8` compile flag; use same heuristics | Both SGEMM and Winograd have INT8 specializations in `*_int8.h` files. |

## Practical Code Examples

### Enabling SGEMM Convolution

Force the matrix-multiplication path for large models:

```cpp
#include "net.h"

int main()
{
    ncnn::Option opt;
    opt.use_sgemm_convolution = true;  // Enable im2col + GEMM
    opt.use_winograd_convolution = false; // Disable Winograd fallback
    opt.num_threads = 4;

    ncnn::Net net;
    net.opt = opt;
    net.load_param("resnet50.param");
    net.load_model("resnet50.bin");
    
    // Inference continues...
}

```

### Enabling Winograd Convolution

Optimize 3×3 layers in ResNet-style architectures:

```cpp
ncnn::Option opt;
opt.use_winograd_convolution = true;  // Enable F(2x2,3) algorithm
opt.use_sgemm_convolution = true;     // Allow SGEMM for non-3x3 layers

ncnn::Net net;
net.opt = opt;
// Load model...

```

### Disabling Optimizations for Debugging

Force the naive implementation to verify numerical correctness:

```cpp
ncnn::Option opt;
opt.use_sgemm_convolution = false;
opt.use_winograd_convolution = false;
// Now uses src/layer/convolution.cpp directly

```

## Key Source Files

| File Path | Description |
|-----------|-------------|
| [`src/layer/convolution.cpp`](https://github.com/Tencent/ncnn/blob/main/src/layer/convolution.cpp) | Baseline direct convolution with nested loops over input, kernel, and output channels. |
| [`src/layer/x86/convolution_x86.cpp`](https://github.com/Tencent/ncnn/blob/main/src/layer/x86/convolution_x86.cpp) | x86 dispatcher that implements `prefer_sgemm` and `prefer_winograd` heuristics to select the optimal path. |
| [`src/layer/x86/convolution_im2col_gemm.h`](https://github.com/Tencent/ncnn/blob/main/src/layer/x86/convolution_im2col_gemm.h) | Im2col packing and kernel transformation for SGEMM-based convolution on x86 (AVX/AVX2/AVX-512). |
| [`src/layer/arm/convolution_im2col_gemm.h`](https://github.com/Tencent/ncnn/blob/main/src/layer/arm/convolution_im2col_gemm.h) | ARM NEON implementation of im2col SGEMM for mobile/embedded devices. |
| [`src/layer/x86/convolution_3x3_winograd.h`](https://github.com/Tencent/ncnn/blob/main/src/layer/x86/convolution_3x3_winograd.h) | Winograd F(2×2, 3) transforms and tile-level computation for 3×3 kernels on x86. |
| [`src/layer/arm/convolution_3x3_winograd.h`](https://github.com/Tencent/ncnn/blob/main/src/layer/arm/convolution_3x3_winograd.h) | ARM NEON Winograd implementation for efficient 3×3 convolution on mobile CPUs. |
| [`src/option.h`](https://github.com/Tencent/ncnn/blob/main/src/option.h) | Runtime configuration struct containing `use_sgemm_convolution` and `use_winograd_convolution` flags. |

## Summary

- **Standard convolution** in [`src/layer/convolution.cpp`](https://github.com/Tencent/ncnn/blob/main/src/layer/convolution.cpp) provides a memory-efficient baseline using direct nested loops, ideal for small tensors or when optimization buffers are prohibitive.
- **SGEMM convolution** transforms inputs via im2col and executes high-performance matrix multiplication, automatically selected when `prefer_sgemm` heuristics detect large channel counts or spatial dimensions that benefit from AVX/AVX-512 or NEON vectorization.
- **Winograd convolution** reduces arithmetic operations by ~2× for 3×3 kernels using the F(2×2, 3) algorithm, activated when `use_winograd_convolution` is enabled and channel counts exceed 8, making it optimal for ResNet-style architectures.

## Frequently Asked Questions

### What is the default convolution implementation in ncnn?

By default, ncnn enables **SGEMM-based convolution** (`opt.use_sgemm_convolution` is true in [`src/option.cpp`](https://github.com/Tencent/ncnn/blob/main/src/option.cpp)). The dispatcher in [`src/layer/x86/convolution_x86.cpp`](https://github.com/Tencent/ncnn/blob/main/src/layer/x86/convolution_x86.cpp) evaluates the `prefer_sgemm` heuristic based on L2 cache size and channel dimensions, falling back to standard convolution only when SGEMM is disabled or suboptimal.

### When should I disable Winograd convolution?

Disable Winograd convolution by setting `opt.use_winograd_convolution = false` when running models with **non-3×3 kernels**, strides greater than 1, or dilation factors other than 1, as the Winograd implementation in [`src/layer/x86/convolution_3x3_winograd.h`](https://github.com/Tencent/ncnn/blob/main/src/layer/x86/convolution_3x3_winograd.h) strictly requires 3×3 kernels with unit stride and dilation. Additionally, disable it for models with fewer than 8 input or output channels where transform overhead exceeds the arithmetic savings.

### How does ncnn decide between SGEMM and Winograd?

The decision logic in [`src/layer/x86/convolution_x86.cpp`](https://github.com/Tencent/ncnn/blob/main/src/layer/x86/convolution_x86.cpp) evaluates boolean flags `prefer_sgemm` and `prefer_winograd`. Winograd is selected when `use_winograd_convolution` is true, the kernel is 3×3 with stride=1 and dilation=1, and `prefer_winograd` (channels ≥ 8) holds true. SGEMM is selected when `use_sgemm_convolution` is true and the tensor size exceeds L2 cache capacity or channels exceed 16, unless Winograd conditions are met and enabled.

### Can I use these optimized convolutions with INT8 quantization?

Yes, both SGEMM and Winograd implementations provide **INT8 specializations** found in corresponding `*_int8.h` files (e.g., [`src/layer/x86/convolution_im2col_gemm_int8.h`](https://github.com/Tencent/ncnn/blob/main/src/layer/x86/convolution_im2col_gemm_int8.h) and [`src/layer/x86/convolution_3x3_winograd_int8.h`](https://github.com/Tencent/ncnn/blob/main/src/layer/x86/convolution_3x3_winograd_int8.h)). To utilize them, compile ncnn with the `NCNN_INT8` macro enabled and apply the same heuristic-based selection via `Option` flags, ensuring that the quantized model weights are properly calibrated for the chosen algorithmic path.