How ncnn Adapts to Different CPU Architectures: ARM, x86, RISC-V, MIPS, and LoongArch Optimization Guide
ncnn detects CPU capabilities at runtime using platform-specific mechanisms in src/cpu.cpp, then automatically dispatches the most optimized SIMD kernels for ARM NEON, x86 AVX-512, RISC-V Vector, and other extensions without requiring recompilation.
Tencent's ncnn is a high-performance neural network inference framework designed for mobile and edge deployment. Understanding ncnn CPU architecture adaptation is essential for maximizing inference speed across diverse hardware, from ARM smartphones to x86 servers and RISC-V IoT devices. The framework achieves this portability through a sophisticated runtime detection layer that queries processor capabilities and selects hand-optimized assembly kernels accordingly.
Runtime CPU Architecture Detection
The foundation of ncnn's portability lies in its CPU abstraction layer, implemented primarily in src/cpu.h and src/cpu.cpp. At library initialization, ncnn queries the host processor's instruction set extensions using operating-system-specific APIs, then caches these results for fast lookup during layer construction.
Platform-Specific Detection Mechanisms
ncnn employs different strategies depending on the target operating system to read CPU capabilities without executing illegal instructions:
- Linux and Android: Reads hardware capability bits from the ELF auxiliary vector using
getauxval(AT_HWCAP)andgetauxval(AT_HWCAP2), falling back to direct/proc/self/auxvparsing when necessary. This exposes ARMHWCAP_NEON,HWCAP_ASIMDHP, and x86HWCAP2_AVX2flags. - macOS: Queries
sysctlbynamefor x86 features (e.g.,hw.optional.avx512f) and ARM64-specific identifiers to determine Apple Silicon capabilities. - Windows: Executes the
CPUIDinstruction via__cpuidand__cpuidexintrinsics, combined with XGETBV register inspection to verify OS-level XSAVE support for AVX and AVX-512 states.
All detection functions follow the naming convention get_cpu_support_{arch}_{feature}() and return integer boolean values.
ARM and AArch64 Feature Detection
For 64-bit ARM (AArch64), ncnn assumes baseline ASIMD (NEON) support, which is mandatory in ARMv8. For 32-bit ARM or extended features, it checks specific hardware capability bits:
int get_cpu_support_arm_neon()
{
#if __aarch64__
return 1; // ASIMD is baseline for aarch64
#else
return (g_hwcaps & HWCAP_NEON) != 0;
#endif
}
The global g_hwcaps variable is populated once at startup from get_elf_hwcap(AT_HWCAP). Advanced extensions like ARM BF16, I8MM, SVE, and SVE2 are detected via HWCAP2 bits in the auxiliary vector, enabling ncnn to leverage scalable vector lengths on server-class ARM processors.
x86 and x86-64 Extension Detection
x86 detection requires multi-leaf CPUID queries. Basic features use leaf 0x01, while modern extensions require leaf 0x07 (sub-leaf 0). Critically, ncnn verifies OS support for extended state management before enabling AVX or AVX-512:
int get_cpu_support_x86_avx()
{
unsigned int cpu_info[4];
x86_cpuid(0, cpu_info);
if (cpu_info[0] < 1) return 0;
x86_cpuid(1, cpu_info);
// Check for AVX bit (28) and OSXSAVE bit (27)
if (!(cpu_info[2] & (1u << 28)) || !(cpu_info[2] & (1u << 27))) return 0;
// Verify XSAVE state is enabled for XMM and YMM registers
if ((x86_get_xcr0() & 0x6) != 0x6) return 0;
return 1;
}
This pattern extends to AVX2, AVX-512 (F, CD, BW, DQ, VL, VNNI, BF16, FP16), and AVX-VNNI for integer operations, ensuring maximum utilization of Intel and AMD server processors.
RISC-V, MIPS, and LoongArch Support
For emerging architectures, ncnn checks single-bit flags in the hardware capability word:
- RISC-V: Detects Vector extension (
COMPAT_HWCAP_ISA_V), half-precision float (Zfh), and the T-Head vector implementation (XTHeadVector) viacpu_support_riscv_v()and related functions. - MIPS: Checks for MSA (MIPS SIMD Architecture) and Loongson-specific MMI extensions.
- LoongArch: Detects LSX (128-bit SIMD) and LASX (256-bit SIMD) vector extensions using architecture-specific hardware capability bits.
Big-Little Core Cluster Management
Modern SoCs combine high-performance "big" cores with power-efficient "little" cores. ncnn distinguishes these clusters to optimize thread scheduling and power consumption during inference.
Heterogeneous CPU Detection
ncnn identifies core types through platform-specific frequency analysis:
- Windows: Reads the
EfficiencyClassfield fromGetLogicalProcessorInformationExsystem calls. - Linux/Android: Measures maximum frequencies via
/sys/devices/system/cpu/cpu*/cpufreq/cpuinfo_max_freqand partitions cores around the median frequency. - macOS: Uses
hw.perflevel0andhw.perflevel1sysctl values to distinguish performance and efficiency cores on Apple Silicon.
The initialization routine initialize_cpu_thread_affinity_mask() (located around line 10,800 in src/cpu.cpp) populates three CpuSet bitmasks:
mask_all: All logical processorsmask_big: High-frequency performance coresmask_little: Low-power efficiency cores
Thread Affinity Control
Users can bind inference threads to specific core types using the public API:
#include "cpu.h"
void configure_for_performance()
{
// Powersave mode 2 = big cores only
const ncnn::CpuSet& big_cores = ncnn::get_cpu_thread_affinity_mask(2);
ncnn::set_cpu_thread_affinity(big_cores);
}
This capability ensures that latency-critical inference runs on big cores while background preprocessing can utilize little cores for battery efficiency.
Kernel Selection and Optimization Strategy
ncnn implements multiple versions of each neural network operator, with file naming conventions indicating the target ISA.
Layer Implementation Variants
Operator implementations follow a predictable pattern in the src/layer/ directory:
conv1x1_fp32_sse.cpp // x86 SSE2 fallback
conv1x1_fp32_avx.cpp // 256-bit AVX
conv1x1_fp32_avx512.cpp // 512-bit AVX-512
conv1x1_fp32_neon.cpp // ARM NEON (128-bit)
conv1x1_fp32_asimdhp.cpp // ARM half-precision (FP16)
conv1x1_fp32_riscv_v.cpp // RISC-V vector extension
conv1x1_fp32_lsx.cpp // LoongArch LSX
Runtime Kernel Dispatch
During layer construction (Layer::create_pipeline()), ncnn queries the cached CPU capability flags and instantiates the most advanced implementation available. The selection logic typically appears as:
if (ncnn::cpu_support_x86_avx512())
pipeline = create_conv1x1_avx512();
else if (ncnn::cpu_support_x86_avx2())
pipeline = create_conv1x1_avx2();
else if (ncnn::cpu_support_arm_neon())
pipeline = create_conv1x1_neon();
else
pipeline = create_conv1x1_fp32(); // Generic C++ fallback
This dispatch mechanism occurs once per layer creation, minimizing overhead while ensuring optimal code paths for the specific silicon.
Practical Implementation Examples
Querying Available Optimizations
Applications can inspect detected features to log or adjust configuration:
#include "cpu.h"
#include <stdio.h>
void print_cpu_capabilities()
{
printf("CPU Count: %d\n", ncnn::get_cpu_count());
if (ncnn::cpu_support_arm_neon())
printf("ARM NEON: available\n");
if (ncnn::cpu_support_x86_avx2())
printf("x86 AVX2: available\n");
if (ncnn::cpu_support_riscv_v())
printf("RISC-V Vector (VLENB=%d): available\n",
ncnn::get_cpu_riscv_vlenb());
}
Pinning Threads to Big Cores
For maximum inference throughput on heterogeneous ARM SoCs:
#include "cpu.h"
void setup_high Performance_inference()
{
// Get mask for big cores (powersave = 2)
const ncnn::CpuSet& big = ncnn::get_cpu_thread_affinity_mask(2);
if (ncnn::set_cpu_thread_affinity(big) == 0)
printf("Successfully pinned to big cores\n");
else
printf("Failed to set thread affinity\n");
}
Cache-Aware Optimization
Custom operators can query cache hierarchy for tiling decisions:
int l2_cache = ncnn::get_cpu_level2_cache_size(); // Bytes
int l3_cache = ncnn::get_cpu_level3_cache_size();
// Adjust tile sizes based on cache capacity
int optimal_tile = (l2_cache / 4) / sizeof(float);
Summary
- ncnn CPU architecture adaptation relies on runtime detection in
src/cpu.cppusinggetauxval(Linux),CPUID(Windows), andsysctl(macOS) to identify available SIMD extensions without recompilation. - Supported architectures include ARM (NEON, SVE, BF16), x86 (AVX, AVX2, AVX-512), RISC-V (Vector, Zfh), MIPS (MSA), and LoongArch (LSX/LASX).
- The framework manages big-little core clusters through
CpuSetmasks accessible viaget_cpu_thread_affinity_mask(), allowing thread pinning to performance or efficiency cores. - Kernel dispatch occurs during layer construction, automatically selecting the most optimized implementation from architecture-specific files like
*_neon.cppor*_avx512.cpp. - Public API functions like
cpu_support_arm_neon()andcpu_support_x86_avx2()enable application-level optimization decisions.
Frequently Asked Questions
How does ncnn detect CPU features without requiring recompilation?
ncnn uses runtime capability detection via operating system APIs. On Linux and Android, it reads the ELF auxiliary vector through getauxval() to check hardware capability bits (e.g., HWCAP_NEON or HWCAP2_AVX2). Windows uses the CPUID instruction combined with XGETBV register checks, while macOS queries sysctlbyname. These mechanisms allow a single binary to run on diverse processors and automatically select appropriate SIMD kernels.
What is the difference between the mask_big and mask_little CpuSets in ncnn?
mask_big represents high-performance CPU cores (higher maximum frequency), while mask_little represents power-efficient cores (lower frequency). ncnn determines these categories by analyzing per-core maximum frequencies on Linux/Android or using the EfficiencyClass field on Windows. Users can retrieve these masks via get_cpu_thread_affinity_mask(2) for big cores or get_cpu_thread_affinity_mask(0) for little cores, then apply them with set_cpu_thread_affinity() to control power consumption and performance.
Does ncnn support AVX-512 on modern x86 processors?
Yes. ncnn detects and utilizes multiple AVX-512 extensions including AVX-512F, CD, BW, DQ, VL, VNNI, BF16, and FP16 through the cpu_support_x86_avx512*() family of functions. The detection verifies both CPU support via CPUID leaf 0x07 and OS support via XCR0 register checks to ensure the processor can save and restore AVX-512 register states during context switches.
How does ncnn handle variable vector lengths on RISC-V processors?
ncnn detects the RISC-V Vector extension through the COMPAT_HWCAP_ISA_V hardware capability bit using cpu_support_riscv_v(). It also queries the specific vector length in bytes (VLENB) via get_cpu_riscv_vlenb() to accommodate processors with different vector register widths. For T-Head implementations, it additionally checks XTHeadVector support, ensuring compatibility with both standard RISC-V vector extensions and vendor-specific variants.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →