Winograd Convolution Optimizations in ncnn: use_winograd23_convolution and use_winograd63_convolution Explained

The use_winograd23_convolution and use_winograd63_convolution flags in Tencent/ncnn enable F(2×2,3×3) and F(6×6,3×3) Winograd algorithms respectively, reducing multiply operations in 3×3 convolutions by up to 4× while trading off transform overhead against data reuse.

Tencent/ncnn provides highly optimized Winograd convolution implementations to accelerate 3×3 kernel operations common in modern CNNs. The library exposes granular control over these optimizations through specific flags in the ncnn::Option structure, allowing developers to tune inference performance based on input dimensions and hardware constraints. Understanding when to enable use_winograd23_convolution versus use_winograd63_convolution is essential for maximizing throughput on both desktop and mobile processors.

Winograd Convolution Variants

ncnn implements two primary Winograd algorithms for 3×3 convolutions with stride 1, distinguished by their tile sizes and computational characteristics.

F(2×2, 3×3) via use_winograd23_convolution

The F(2×2, 3×3) variant processes 2×2 output tiles, reducing the multiplication count from 9 to 4 per tile. This smaller tile size minimizes transform overhead and temporary buffer requirements, making it suitable for medium spatial dimensions (approximately 8–30 pixels) and moderate channel counts (≤32–64). In src/option.cpp, this optimization defaults to enabled:

use_winograd23_convolution = true;   // src/option.cpp#L64

F(6×6, 3×3) via use_winograd63_convolution

The F(6×6, 3×3) variant uses larger 6×6 output tiles, achieving the same theoretical arithmetic reduction (9→4 operations) but with significantly better data reuse. This amortizes the transform cost across more elements, yielding higher throughput for large spatial sizes (≥30 pixels) on modern CPUs with sufficient cache. The flag is defined in the same file:

use_winograd63_convolution = true;   // src/option.cpp#L66

Runtime Selection Strategy

ncnn employs architecture-specific heuristics to automatically select between Winograd variants during layer creation. The selection logic resides in backend-specific convolution implementations and evaluates input/output channel counts and spatial dimensions against empirical thresholds.

On x86 platforms, the decision functions test_prefer_winograd63 and test_prefer_winograd23 in src/layer/x86/convolution_x86.cpp implement these heuristics. The winograd63 test occupies lines 105–141, while winograd23 logic appears at lines 175–210. Similar selection strategies exist for ARM (src/layer/arm/convolution_arm.cpp), RISC-V (src/layer/riscv/convolution_riscv.cpp), MIPS, and LoongArch backends, each optimized for their respective microarchitectures.

The selection flow follows this pattern:

bool prefer_winograd = (opt.use_winograd23_convolution ||
                        opt.use_winograd43_convolution ||
                        opt.use_winograd63_convolution) &&
                       (num_input > 8 || num_output > 8);

if (opt.use_winograd_convolution && prefer_winograd &&
    kernel_w == 3 && kernel_h == 3 && stride_w == 1 && stride_h == 1) {
    bool prefer_winograd63 = test_prefer_winograd63(num_input, num_output, w, h);
    bool prefer_winograd23 = test_prefer_winograd23(num_input, num_output, w, h);
    // Variant selection with fallback logic
}

Use Cases and Performance Tuning

Choosing between these flags depends on feature map dimensions, memory constraints, and target hardware.

Default Configuration: Keep both flags true to allow the runtime automatic selection. This provides optimal performance for most models on modern CPUs.

Small Feature Maps: Disable use_winograd63_convolution when processing very small inputs (e.g., 4×4 feature maps). The transform overhead for 6×6 tiles outweighs the computational savings at small spatial dimensions.

Memory-Constrained Environments: On embedded devices with limited RAM (e.g., ARM Cortex-A53), disable use_winograd63_convolution while keeping use_winograd23_convolution enabled. The 2×2 tile variant requires less temporary storage for intermediate transforms.

Integer Quantization: When running INT8 inference, specialized Winograd kernels (e.g., utilizing AVX-512-VNNI on x86) may be unavailable. Disabling the specific flags prevents fallback to suboptimal SGEMM paths when the quantized Winograd implementations are missing.

Implementation Details

Each Winograd variant maintains separate transform and computation pipelines.

Weight Transformation: The functions conv3x3s1_winograd23_transform_kernel and conv3x3s1_winograd63_transform_kernel pre-transform convolution weights during model loading. These transformed weights are stored in weight_winograd23_data and weight_winograd63_data respectively, avoiding redundant transforms during inference.

Computation Kernels: The actual tile computations are implemented in conv3x3s1_winograd23 and conv3x3s1_winograd63, with additional variants supporting INT8, FP16, and BF16 data types.

Fallback Behavior: If a specific variant is disabled (e.g., use_winograd23_convolution == false), the selection logic attempts the next available variant (typically falling back from 63 to 23 to 43, or ultimately to direct SGEMM convolution).

Configuration Examples

The following examples demonstrate how to configure these options in C++ and Python.

C++ Configuration

#include "net.h"

int main() {
    ncnn::Net net;
    net.load_param("model.param");
    net.load_model("model.bin");

    ncnn::Option opt = net.opt;
    opt.use_winograd_convolution = true;      // Master switch
    opt.use_winograd23_convolution = true;    // Enable 2×2 tiles
    opt.use_winograd63_convolution = false;   // Disable 6×6 tiles for small inputs
    
    net.opt = opt;
    
    // Inference continues...
}

Python Configuration

import ncnn

net = ncnn.Net()
net.load_param("model.param")
net.load_model("model.bin")

opt = net.opt
opt.use_winograd_convolution = True
opt.use_winograd23_convolution = True
opt.use_winograd63_convolution = False  # Disable for memory-constrained inference

net.opt = opt

Summary

  • Two variants: use_winograd23_convolution enables F(2×2,3×3) for smaller tiles and lower memory usage, while use_winograd63_convolution enables F(6×6,3×3) for better data reuse on large feature maps.
  • Automatic selection: ncnn uses heuristic functions (test_prefer_winograd23, test_prefer_winograd63) in architecture-specific files like src/layer/x86/convolution_x86.cpp to choose the optimal variant at runtime.
  • Configuration: Both flags default to true in src/option.cpp, but can be disabled individually to reduce memory footprint on embedded devices or avoid transform overhead on small inputs.
  • Implementation: Each variant uses dedicated weight transform functions and data storage buffers, with automatic fallback to direct convolution if Winograd is disabled or unsuitable.

Frequently Asked Questions

What is the difference between use_winograd23_convolution and use_winograd63_convolution?

use_winograd23_convolution enables the F(2×2,3×3) algorithm using 2×2 output tiles, optimal for medium-sized feature maps (8–30 pixels) and memory-constrained environments. use_winograd63_convolution enables F(6×6,3×3) with 6×6 tiles, providing superior data reuse and throughput for large spatial dimensions (≥30 pixels) at the cost of increased temporary buffer requirements.

When should I disable these Winograd optimizations?

Disable use_winograd63_convolution for very small feature maps (under 8×8) where transform overhead exceeds computational savings, or on memory-constrained embedded devices where the 6×6 tile buffer allocation is prohibitive. Disable both flags when debugging numerical accuracy issues to force standard SGEMM convolution and isolate precision differences.

Does the selection logic differ between CPU architectures?

Yes. While the flags are defined globally in ncnn::Option, the selection heuristics in test_prefer_winograd23 and test_prefer_winograd63 are implemented separately for each backend (x86, ARM, RISC-V, MIPS, LoongArch) in their respective convolution_[arch].cpp files. Thresholds vary to match each architecture's cache size, SIMD capabilities, and memory bandwidth characteristics.

Can these flags be used with INT8 quantized models?

Yes. ncnn provides INT8-optimized Winograd kernels for both variants on supported architectures (e.g., ARM NEON and x86 AVX-512-VNNI). If the specific INT8 Winograd implementation is unavailable for your target, the flags automatically prevent selection of that path, falling back to INT8 SGEMM or direct convolution to maintain correctness.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →