# Winograd Convolution Optimizations in ncnn: use_winograd23_convolution and use_winograd63_convolution Explained

> Explore Winograd convolution optimizations in ncnn with use_winograd23_convolution and use_winograd63_convolution. Reduce multiply operations in 3x3 convolutions and boost performance.

- Repository: [Tencent/ncnn](https://github.com/tencent/ncnn)
- Tags: deep-dive
- Published: 2026-02-23

---

**The `use_winograd23_convolution` and `use_winograd63_convolution` flags in Tencent/ncnn enable F(2×2,3×3) and F(6×6,3×3) Winograd algorithms respectively, reducing multiply operations in 3×3 convolutions by up to 4× while trading off transform overhead against data reuse.**

Tencent/ncnn provides highly optimized Winograd convolution implementations to accelerate 3×3 kernel operations common in modern CNNs. The library exposes granular control over these optimizations through specific flags in the `ncnn::Option` structure, allowing developers to tune inference performance based on input dimensions and hardware constraints. Understanding when to enable `use_winograd23_convolution` versus `use_winograd63_convolution` is essential for maximizing throughput on both desktop and mobile processors.

## Winograd Convolution Variants

ncnn implements two primary Winograd algorithms for 3×3 convolutions with stride 1, distinguished by their tile sizes and computational characteristics.

### F(2×2, 3×3) via use_winograd23_convolution

The **F(2×2, 3×3)** variant processes 2×2 output tiles, reducing the multiplication count from 9 to 4 per tile. This smaller tile size minimizes transform overhead and temporary buffer requirements, making it suitable for medium spatial dimensions (approximately 8–30 pixels) and moderate channel counts (≤32–64). In [`src/option.cpp`](https://github.com/Tencent/ncnn/blob/main/src/option.cpp), this optimization defaults to enabled:

```cpp
use_winograd23_convolution = true;   // src/option.cpp#L64

```

### F(6×6, 3×3) via use_winograd63_convolution

The **F(6×6, 3×3)** variant uses larger 6×6 output tiles, achieving the same theoretical arithmetic reduction (9→4 operations) but with significantly better data reuse. This amortizes the transform cost across more elements, yielding higher throughput for large spatial sizes (≥30 pixels) on modern CPUs with sufficient cache. The flag is defined in the same file:

```cpp
use_winograd63_convolution = true;   // src/option.cpp#L66

```

## Runtime Selection Strategy

ncnn employs architecture-specific heuristics to automatically select between Winograd variants during layer creation. The selection logic resides in backend-specific convolution implementations and evaluates input/output channel counts and spatial dimensions against empirical thresholds.

On x86 platforms, the decision functions `test_prefer_winograd63` and `test_prefer_winograd23` in [`src/layer/x86/convolution_x86.cpp`](https://github.com/Tencent/ncnn/blob/main/src/layer/x86/convolution_x86.cpp) implement these heuristics. The winograd63 test occupies lines 105–141, while winograd23 logic appears at lines 175–210. Similar selection strategies exist for ARM ([`src/layer/arm/convolution_arm.cpp`](https://github.com/Tencent/ncnn/blob/main/src/layer/arm/convolution_arm.cpp)), RISC-V ([`src/layer/riscv/convolution_riscv.cpp`](https://github.com/Tencent/ncnn/blob/main/src/layer/riscv/convolution_riscv.cpp)), MIPS, and LoongArch backends, each optimized for their respective microarchitectures.

The selection flow follows this pattern:

```cpp
bool prefer_winograd = (opt.use_winograd23_convolution ||
                        opt.use_winograd43_convolution ||
                        opt.use_winograd63_convolution) &&
                       (num_input > 8 || num_output > 8);

if (opt.use_winograd_convolution && prefer_winograd &&
    kernel_w == 3 && kernel_h == 3 && stride_w == 1 && stride_h == 1) {
    bool prefer_winograd63 = test_prefer_winograd63(num_input, num_output, w, h);
    bool prefer_winograd23 = test_prefer_winograd23(num_input, num_output, w, h);
    // Variant selection with fallback logic
}

```

## Use Cases and Performance Tuning

Choosing between these flags depends on feature map dimensions, memory constraints, and target hardware.

**Default Configuration:** Keep both flags `true` to allow the runtime automatic selection. This provides optimal performance for most models on modern CPUs.

**Small Feature Maps:** Disable `use_winograd63_convolution` when processing very small inputs (e.g., 4×4 feature maps). The transform overhead for 6×6 tiles outweighs the computational savings at small spatial dimensions.

**Memory-Constrained Environments:** On embedded devices with limited RAM (e.g., ARM Cortex-A53), disable `use_winograd63_convolution` while keeping `use_winograd23_convolution` enabled. The 2×2 tile variant requires less temporary storage for intermediate transforms.

**Integer Quantization:** When running INT8 inference, specialized Winograd kernels (e.g., utilizing AVX-512-VNNI on x86) may be unavailable. Disabling the specific flags prevents fallback to suboptimal SGEMM paths when the quantized Winograd implementations are missing.

## Implementation Details

Each Winograd variant maintains separate transform and computation pipelines.

**Weight Transformation:** The functions `conv3x3s1_winograd23_transform_kernel` and `conv3x3s1_winograd63_transform_kernel` pre-transform convolution weights during model loading. These transformed weights are stored in `weight_winograd23_data` and `weight_winograd63_data` respectively, avoiding redundant transforms during inference.

**Computation Kernels:** The actual tile computations are implemented in `conv3x3s1_winograd23` and `conv3x3s1_winograd63`, with additional variants supporting INT8, FP16, and BF16 data types.

**Fallback Behavior:** If a specific variant is disabled (e.g., `use_winograd23_convolution == false`), the selection logic attempts the next available variant (typically falling back from 63 to 23 to 43, or ultimately to direct SGEMM convolution).

## Configuration Examples

The following examples demonstrate how to configure these options in C++ and Python.

### C++ Configuration

```cpp
#include "net.h"

int main() {
    ncnn::Net net;
    net.load_param("model.param");
    net.load_model("model.bin");

    ncnn::Option opt = net.opt;
    opt.use_winograd_convolution = true;      // Master switch
    opt.use_winograd23_convolution = true;    // Enable 2×2 tiles
    opt.use_winograd63_convolution = false;   // Disable 6×6 tiles for small inputs
    
    net.opt = opt;
    
    // Inference continues...
}

```

### Python Configuration

```python
import ncnn

net = ncnn.Net()
net.load_param("model.param")
net.load_model("model.bin")

opt = net.opt
opt.use_winograd_convolution = True
opt.use_winograd23_convolution = True
opt.use_winograd63_convolution = False  # Disable for memory-constrained inference

net.opt = opt

```

## Summary

- **Two variants:** `use_winograd23_convolution` enables F(2×2,3×3) for smaller tiles and lower memory usage, while `use_winograd63_convolution` enables F(6×6,3×3) for better data reuse on large feature maps.
- **Automatic selection:** ncnn uses heuristic functions (`test_prefer_winograd23`, `test_prefer_winograd63`) in architecture-specific files like [`src/layer/x86/convolution_x86.cpp`](https://github.com/Tencent/ncnn/blob/main/src/layer/x86/convolution_x86.cpp) to choose the optimal variant at runtime.
- **Configuration:** Both flags default to `true` in [`src/option.cpp`](https://github.com/Tencent/ncnn/blob/main/src/option.cpp), but can be disabled individually to reduce memory footprint on embedded devices or avoid transform overhead on small inputs.
- **Implementation:** Each variant uses dedicated weight transform functions and data storage buffers, with automatic fallback to direct convolution if Winograd is disabled or unsuitable.

## Frequently Asked Questions

### What is the difference between use_winograd23_convolution and use_winograd63_convolution?

`use_winograd23_convolution` enables the F(2×2,3×3) algorithm using 2×2 output tiles, optimal for medium-sized feature maps (8–30 pixels) and memory-constrained environments. `use_winograd63_convolution` enables F(6×6,3×3) with 6×6 tiles, providing superior data reuse and throughput for large spatial dimensions (≥30 pixels) at the cost of increased temporary buffer requirements.

### When should I disable these Winograd optimizations?

Disable `use_winograd63_convolution` for very small feature maps (under 8×8) where transform overhead exceeds computational savings, or on memory-constrained embedded devices where the 6×6 tile buffer allocation is prohibitive. Disable both flags when debugging numerical accuracy issues to force standard SGEMM convolution and isolate precision differences.

### Does the selection logic differ between CPU architectures?

Yes. While the flags are defined globally in `ncnn::Option`, the selection heuristics in `test_prefer_winograd23` and `test_prefer_winograd63` are implemented separately for each backend (x86, ARM, RISC-V, MIPS, LoongArch) in their respective `convolution_[arch].cpp` files. Thresholds vary to match each architecture's cache size, SIMD capabilities, and memory bandwidth characteristics.

### Can these flags be used with INT8 quantized models?

Yes. ncnn provides INT8-optimized Winograd kernels for both variants on supported architectures (e.g., ARM NEON and x86 AVX-512-VNNI). If the specific INT8 Winograd implementation is unavailable for your target, the flags automatically prevent selection of that path, falling back to INT8 SGEMM or direct convolution to maintain correctness.