# How ncnn Adapts to Different CPU Architectures: ARM, x86, RISC-V, MIPS, and LoongArch Optimization Guide

> Explore ncnn's CPU architecture adaptation for ARM, x86, RISC-V, MIPS, and LoongArch. Discover runtime detection and auto-optimization for peak performance without recompilation.

- Repository: [Tencent/ncnn](https://github.com/tencent/ncnn)
- Tags: deep-dive
- Published: 2026-02-23

---

**ncnn detects CPU capabilities at runtime using platform-specific mechanisms in [`src/cpu.cpp`](https://github.com/Tencent/ncnn/blob/main/src/cpu.cpp), then automatically dispatches the most optimized SIMD kernels for ARM NEON, x86 AVX-512, RISC-V Vector, and other extensions without requiring recompilation.**

Tencent's ncnn is a high-performance neural network inference framework designed for mobile and edge deployment. Understanding **ncnn CPU architecture adaptation** is essential for maximizing inference speed across diverse hardware, from ARM smartphones to x86 servers and RISC-V IoT devices. The framework achieves this portability through a sophisticated runtime detection layer that queries processor capabilities and selects hand-optimized assembly kernels accordingly.

## Runtime CPU Architecture Detection

The foundation of ncnn's portability lies in its CPU abstraction layer, implemented primarily in [`src/cpu.h`](https://github.com/Tencent/ncnn/blob/main/src/cpu.h) and [`src/cpu.cpp`](https://github.com/Tencent/ncnn/blob/main/src/cpu.cpp). At library initialization, ncnn queries the host processor's instruction set extensions using operating-system-specific APIs, then caches these results for fast lookup during layer construction.

### Platform-Specific Detection Mechanisms

ncnn employs different strategies depending on the target operating system to read CPU capabilities without executing illegal instructions:

- **Linux and Android**: Reads hardware capability bits from the ELF auxiliary vector using `getauxval(AT_HWCAP)` and `getauxval(AT_HWCAP2)`, falling back to direct `/proc/self/auxv` parsing when necessary. This exposes ARM `HWCAP_NEON`, `HWCAP_ASIMDHP`, and x86 `HWCAP2_AVX2` flags.
- **macOS**: Queries `sysctlbyname` for x86 features (e.g., `hw.optional.avx512f`) and ARM64-specific identifiers to determine Apple Silicon capabilities.
- **Windows**: Executes the `CPUID` instruction via `__cpuid` and `__cpuidex` intrinsics, combined with XGETBV register inspection to verify OS-level XSAVE support for AVX and AVX-512 states.

All detection functions follow the naming convention `get_cpu_support_{arch}_{feature}()` and return integer boolean values.

### ARM and AArch64 Feature Detection

For 64-bit ARM (AArch64), ncnn assumes baseline ASIMD (NEON) support, which is mandatory in ARMv8. For 32-bit ARM or extended features, it checks specific hardware capability bits:

```cpp
int get_cpu_support_arm_neon()
{
#if __aarch64__
    return 1;  // ASIMD is baseline for aarch64
#else
    return (g_hwcaps & HWCAP_NEON) != 0;
#endif
}

```

The global `g_hwcaps` variable is populated once at startup from `get_elf_hwcap(AT_HWCAP)`. Advanced extensions like **ARM BF16**, **I8MM**, **SVE**, and **SVE2** are detected via `HWCAP2` bits in the auxiliary vector, enabling ncnn to leverage scalable vector lengths on server-class ARM processors.

### x86 and x86-64 Extension Detection

x86 detection requires multi-leaf CPUID queries. Basic features use leaf `0x01`, while modern extensions require leaf `0x07` (sub-leaf 0). Critically, ncnn verifies OS support for extended state management before enabling AVX or AVX-512:

```cpp
int get_cpu_support_x86_avx()
{
    unsigned int cpu_info[4];
    x86_cpuid(0, cpu_info);
    if (cpu_info[0] < 1) return 0;

    x86_cpuid(1, cpu_info);
    // Check for AVX bit (28) and OSXSAVE bit (27)
    if (!(cpu_info[2] & (1u << 28)) || !(cpu_info[2] & (1u << 27))) return 0;
    
    // Verify XSAVE state is enabled for XMM and YMM registers
    if ((x86_get_xcr0() & 0x6) != 0x6) return 0;
    
    return 1;
}

```

This pattern extends to **AVX2**, **AVX-512** (F, CD, BW, DQ, VL, VNNI, BF16, FP16), and **AVX-VNNI** for integer operations, ensuring maximum utilization of Intel and AMD server processors.

### RISC-V, MIPS, and LoongArch Support

For emerging architectures, ncnn checks single-bit flags in the hardware capability word:

- **RISC-V**: Detects Vector extension (`COMPAT_HWCAP_ISA_V`), half-precision float (`Zfh`), and the T-Head vector implementation (`XTHeadVector`) via `cpu_support_riscv_v()` and related functions.
- **MIPS**: Checks for **MSA** (MIPS SIMD Architecture) and Loongson-specific **MMI** extensions.
- **LoongArch**: Detects **LSX** (128-bit SIMD) and **LASX** (256-bit SIMD) vector extensions using architecture-specific hardware capability bits.

## Big-Little Core Cluster Management

Modern SoCs combine high-performance "big" cores with power-efficient "little" cores. ncnn distinguishes these clusters to optimize thread scheduling and power consumption during inference.

### Heterogeneous CPU Detection

ncnn identifies core types through platform-specific frequency analysis:

- **Windows**: Reads the `EfficiencyClass` field from `GetLogicalProcessorInformationEx` system calls.
- **Linux/Android**: Measures maximum frequencies via `/sys/devices/system/cpu/cpu*/cpufreq/cpuinfo_max_freq` and partitions cores around the median frequency.
- **macOS**: Uses `hw.perflevel0` and `hw.perflevel1` sysctl values to distinguish performance and efficiency cores on Apple Silicon.

The initialization routine `initialize_cpu_thread_affinity_mask()` (located around line 10,800 in [`src/cpu.cpp`](https://github.com/Tencent/ncnn/blob/main/src/cpu.cpp)) populates three `CpuSet` bitmasks:

- `mask_all`: All logical processors
- `mask_big`: High-frequency performance cores
- `mask_little`: Low-power efficiency cores

### Thread Affinity Control

Users can bind inference threads to specific core types using the public API:

```cpp
#include "cpu.h"

void configure_for_performance()
{
    // Powersave mode 2 = big cores only
    const ncnn::CpuSet& big_cores = ncnn::get_cpu_thread_affinity_mask(2);
    ncnn::set_cpu_thread_affinity(big_cores);
}

```

This capability ensures that latency-critical inference runs on big cores while background preprocessing can utilize little cores for battery efficiency.

## Kernel Selection and Optimization Strategy

ncnn implements multiple versions of each neural network operator, with file naming conventions indicating the target ISA.

### Layer Implementation Variants

Operator implementations follow a predictable pattern in the `src/layer/` directory:

```

conv1x1_fp32_sse.cpp       // x86 SSE2 fallback
conv1x1_fp32_avx.cpp       // 256-bit AVX
conv1x1_fp32_avx512.cpp    // 512-bit AVX-512
conv1x1_fp32_neon.cpp      // ARM NEON (128-bit)
conv1x1_fp32_asimdhp.cpp   // ARM half-precision (FP16)
conv1x1_fp32_riscv_v.cpp   // RISC-V vector extension
conv1x1_fp32_lsx.cpp       // LoongArch LSX

```

### Runtime Kernel Dispatch

During layer construction (`Layer::create_pipeline()`), ncnn queries the cached CPU capability flags and instantiates the most advanced implementation available. The selection logic typically appears as:

```cpp
if (ncnn::cpu_support_x86_avx512())
    pipeline = create_conv1x1_avx512();
else if (ncnn::cpu_support_x86_avx2())
    pipeline = create_conv1x1_avx2();
else if (ncnn::cpu_support_arm_neon())
    pipeline = create_conv1x1_neon();
else
    pipeline = create_conv1x1_fp32();  // Generic C++ fallback

```

This dispatch mechanism occurs once per layer creation, minimizing overhead while ensuring optimal code paths for the specific silicon.

## Practical Implementation Examples

### Querying Available Optimizations

Applications can inspect detected features to log or adjust configuration:

```cpp
#include "cpu.h"
#include <stdio.h>

void print_cpu_capabilities()
{
    printf("CPU Count: %d\n", ncnn::get_cpu_count());
    
    if (ncnn::cpu_support_arm_neon())
        printf("ARM NEON: available\n");
        
    if (ncnn::cpu_support_x86_avx2())
        printf("x86 AVX2: available\n");
        
    if (ncnn::cpu_support_riscv_v())
        printf("RISC-V Vector (VLENB=%d): available\n", 
               ncnn::get_cpu_riscv_vlenb());
}

```

### Pinning Threads to Big Cores

For maximum inference throughput on heterogeneous ARM SoCs:

```cpp
#include "cpu.h"

void setup_high Performance_inference()
{
    // Get mask for big cores (powersave = 2)
    const ncnn::CpuSet& big = ncnn::get_cpu_thread_affinity_mask(2);
    
    if (ncnn::set_cpu_thread_affinity(big) == 0)
        printf("Successfully pinned to big cores\n");
    else
        printf("Failed to set thread affinity\n");
}

```

### Cache-Aware Optimization

Custom operators can query cache hierarchy for tiling decisions:

```cpp
int l2_cache = ncnn::get_cpu_level2_cache_size();  // Bytes
int l3_cache = ncnn::get_cpu_level3_cache_size();

// Adjust tile sizes based on cache capacity
int optimal_tile = (l2_cache / 4) / sizeof(float);

```

## Summary

- **ncnn CPU architecture adaptation** relies on runtime detection in [`src/cpu.cpp`](https://github.com/Tencent/ncnn/blob/main/src/cpu.cpp) using `getauxval` (Linux), `CPUID` (Windows), and `sysctl` (macOS) to identify available SIMD extensions without recompilation.
- Supported architectures include **ARM** (NEON, SVE, BF16), **x86** (AVX, AVX2, AVX-512), **RISC-V** (Vector, Zfh), **MIPS** (MSA), and **LoongArch** (LSX/LASX).
- The framework manages **big-little core clusters** through `CpuSet` masks accessible via `get_cpu_thread_affinity_mask()`, allowing thread pinning to performance or efficiency cores.
- Kernel dispatch occurs during layer construction, automatically selecting the most optimized implementation from architecture-specific files like `*_neon.cpp` or `*_avx512.cpp`.
- Public API functions like `cpu_support_arm_neon()` and `cpu_support_x86_avx2()` enable application-level optimization decisions.

## Frequently Asked Questions

### How does ncnn detect CPU features without requiring recompilation?

ncnn uses **runtime capability detection** via operating system APIs. On Linux and Android, it reads the ELF auxiliary vector through `getauxval()` to check hardware capability bits (e.g., `HWCAP_NEON` or `HWCAP2_AVX2`). Windows uses the `CPUID` instruction combined with XGETBV register checks, while macOS queries `sysctlbyname`. These mechanisms allow a single binary to run on diverse processors and automatically select appropriate SIMD kernels.

### What is the difference between the mask_big and mask_little CpuSets in ncnn?

`mask_big` represents high-performance CPU cores (higher maximum frequency), while `mask_little` represents power-efficient cores (lower frequency). ncnn determines these categories by analyzing per-core maximum frequencies on Linux/Android or using the `EfficiencyClass` field on Windows. Users can retrieve these masks via `get_cpu_thread_affinity_mask(2)` for big cores or `get_cpu_thread_affinity_mask(0)` for little cores, then apply them with `set_cpu_thread_affinity()` to control power consumption and performance.

### Does ncnn support AVX-512 on modern x86 processors?

Yes. ncnn detects and utilizes multiple AVX-512 extensions including **AVX-512F**, **CD**, **BW**, **DQ**, **VL**, **VNNI**, **BF16**, and **FP16** through the `cpu_support_x86_avx512*()` family of functions. The detection verifies both CPU support via CPUID leaf 0x07 and OS support via XCR0 register checks to ensure the processor can save and restore AVX-512 register states during context switches.

### How does ncnn handle variable vector lengths on RISC-V processors?

ncnn detects the RISC-V Vector extension through the `COMPAT_HWCAP_ISA_V` hardware capability bit using `cpu_support_riscv_v()`. It also queries the specific vector length in bytes (VLENB) via `get_cpu_riscv_vlenb()` to accommodate processors with different vector register widths. For T-Head implementations, it additionally checks `XTHeadVector` support, ensuring compatibility with both standard RISC-V vector extensions and vendor-specific variants.