# CPU Architectures Supported by BitNet: x86-64 and ARM64 Optimization Guide

> Discover BitNet's CPU architecture support for x86-64 and ARM64. Learn about AVX2, AVX-512, NEON, and DOT-PROD optimizations tailored for each.

- Repository: [Microsoft/BitNet](https://github.com/microsoft/BitNet)
- Tags: optimization-guide
- Published: 2026-03-13

---

**BitNet supports x86-64 (Intel/AMD) and ARM64 CPUs, leveraging AVX2/AVX-512 for x86 and NEON/DOT-PROD for ARM through architecture-specific TL2 and TL1 kernels.**

The microsoft/BitNet repository delivers a high-performance inference engine optimized for modern CPU architectures. Understanding the CPU architectures supported by BitNet is essential for maximizing throughput, as the project implements distinct kernel variants and tiling strategies tailored to x86-64 and ARM64 instruction sets.

## BitNet CPU Architecture Support Matrix

BitNet’s inference engine targets two primary processor families with specialized code paths:

| Architecture | Instruction Sets | Kernel Variant | Key Tiling Parameters |
|--------------|------------------|----------------|----------------------|
| **x86-64** (Intel/AMD) | AVX, AVX2, AVX-512F, SSSE3 | **TL2** ([`bitnet-lut-kernels-tl2.h`](https://github.com/microsoft/BitNet/blob/main/bitnet-lut-kernels-tl2.h)) | Row-block: 4, Col-block: 128, Parallel: 4 |
| **ARM64** (aarch64) | NEON, DOT-PROD (optional) | **TL1** ([`bitnet-lut-kernels-tl1.h`](https://github.com/microsoft/BitNet/blob/main/bitnet-lut-kernels-tl1.h)) | Row-block: 8, Col-block: 256, Parallel: 8 (with DOT-PROD) |

The repository explicitly documents these targets in [`src/README.md`](https://github.com/microsoft/BitNet/blob/main/src/README.md), noting support for x86-64 with AVX2, ARM with NEON, and ARM with DOTPROD extensions.

## x86-64 Optimizations in BitNet

### Instruction Set Features (AVX2, AVX-512, SSSE3)

The x86-64 implementation detects available vector extensions at compile time using preprocessor macros. In [`include/gemm-config.h`](https://github.com/microsoft/BitNet/blob/main/include/gemm-config.h) and kernel headers, BitNet checks for `__AVX__`, `__AVX2__`, `__AVX512F__`, and `__SSSE3__` to enable vectorized paths.

The primary computation uses **W2A8** (weight-2-bit, activation-8-bit) `vet_dot` kernels that pack weights and activations into 256-bit (AVX2) or 512-bit (AVX-512) registers for parallel dot-product accumulation.

### TL2 Kernel Implementation

x86-64 platforms use the **TL2** kernel variant defined in [`bitnet-lut-kernels-tl2.h`](https://github.com/microsoft/BitNet/blob/main/bitnet-lut-kernels-tl2.h). These kernels implement lookup-table-based quantization and activation-parallel processing strategies optimized for the x86 memory hierarchy.

The implementation resides in [`src/ggml-bitnet-mad.cpp`](https://github.com/microsoft/BitNet/blob/main/src/ggml-bitnet-mad.cpp), which contains the parallel W2A8 kernels for both architectures.

### Tiling Parameters for x86-64

BitNet uses cache-friendly GEMM tiling configured in [`include/gemm-config.h`](https://github.com/microsoft/BitNet/blob/main/include/gemm-config.h). For x86-64 with `ACT_PARALLEL` defined, the defaults are:

```c
#define ROW_BLOCK_SIZE 4
#define COL_BLOCK_SIZE 128
#define PARALLEL_SIZE 4

```

These values optimize for L1/L2 cache utilization on modern Intel and AMD processors when processing 4 rows simultaneously with 128-column blocks.

To build with x86-64 TL2 optimizations:

```bash
cmake -B build -DCMAKE_BUILD_TYPE=Release -DBITNET_X86_TL2=ON
cmake --build build -j$(nproc)

```

## ARM64 Optimizations in BitNet

### NEON and DOT-PROD Extensions

The ARM64 implementation targets AArch64 processors with **NEON** SIMD support, detected via `__ARM_NEON`. For additional performance, BitNet leverages the **DOT-PROD** extension (`__ARM_FEATURE_DOTPROD`), which provides specialized instructions for 8-bit dot-product accumulation critical for the W2A8 kernels.

### TL1 Kernel Implementation

ARM64 platforms use the **TL1** kernel variant defined in [`bitnet-lut-kernels-tl1.h`](https://github.com/microsoft/BitNet/blob/main/bitnet-lut-kernels-tl1.h). These kernels implement activation-parallel strategies using NEON registers, with optimized paths when DOT-PROD instructions are available.

The kernel selection occurs at compile time based on the `BITNET_ARM_TL1` CMake option, which defines `GGML_BITNET_ARM_TL1` and includes the TL1 headers.

### Tiling Parameters for ARM64

In [`include/gemm-config.h`](https://github.com/microsoft/BitNet/blob/main/include/gemm-config.h), ARM64 uses different tiling defaults optimized for the ARM memory hierarchy. When `__ARM_FEATURE_DOTPROD` is present with `ACT_PARALLEL`:

```c
#define ROW_BLOCK_SIZE 8
#define COL_BLOCK_SIZE 256
#define PARALLEL_SIZE 8

```

Without DOT-PROD, the configuration uses more conservative tiling to accommodate the standard NEON instruction latency.

To build for ARM64 (including Apple Silicon):

```bash
cmake -B build -DCMAKE_BUILD_TYPE=Release -DBITNET_ARM_TL1=ON
cmake --build build -j$(sysctl -n hw.ncpu)

```

## Build Configuration and Architecture Detection

### CMake Options for Architecture Selection

The build system exposes explicit options in [`CMakeLists.txt`](https://github.com/microsoft/BitNet/blob/main/CMakeLists.txt) (lines 15-33) to select kernel variants:

- `BITNET_X86_TL2`: Enables x86-64 TL2 kernels, defining `GGML_BITNET_X86_TL2`
- `BITNET_ARM_TL1`: Enables ARM64 TL1 kernels, defining `GGML_BITNET_ARM_TL1`

These flags determine which kernel headers are included and which code paths are compiled into the final binary.

### Automatic Detection via setup_env.py

The [`setup_env.py`](https://github.com/microsoft/BitNet/blob/main/setup_env.py) script automates architecture detection using `platform.machine()` to identify `x86_64` or `aarch64` hosts. Based on detection, it injects appropriate compiler flags via the `COMPILER_EXTRA_ARGS` dictionary (lines 66-69):

- For x86_64: Sets quantization type to `i2_s` and enables TL2 kernels
- For ARM64: Sets quantization type to `tl1` or `tl2` and enables TL1 kernels

To use the automatic configuration:

```bash
python setup_env.py \
    --model-dir ./models/BitNet-b1.58-2B-4T \
    --quant-type tl2 \
    --use-pretuned

```

The script also copies pre-tuned kernel configurations from `preset_kernels/` based on the detected architecture.

## Customizing Tiling for Specific CPUs

For specialized deployments, you can override the default tiling parameters in [`include/gemm-config.h`](https://github.com/microsoft/BitNet/blob/main/include/gemm-config.h) to match your processor's cache hierarchy. For example, to optimize for a larger L2 cache on x86:

```c
/* Custom tiling for large L2 cache x86 processors */
#if defined(__AVX2__) && defined(ACT_PARALLEL)

#   define ROW_BLOCK_SIZE 8

#   define COL_BLOCK_SIZE 256

#   define PARALLEL_SIZE 8

#endif

```

After modifying the header, rebuild the project:

```bash
cmake --build build -j$(nproc)

```

## Summary

- BitNet supports **x86-64** (Intel/AMD) and **ARM64** (including Apple Silicon) through specialized kernel implementations.
- **x86-64** uses **TL2 kernels** with AVX2/AVX-512 instruction sets, configured with 4×128 tiling parameters.
- **ARM64** uses **TL1 kernels** with NEON and optional DOT-PROD extensions, configured with 8×256 tiling when DOT-PROD is available.
- The build system auto-detects architecture via [`setup_env.py`](https://github.com/microsoft/BitNet/blob/main/setup_env.py) and exposes CMake options `BITNET_X86_TL2` and `BITNET_ARM_TL1` for manual selection.
- Tuning parameters are centralized in [`include/gemm-config.h`](https://github.com/microsoft/BitNet/blob/main/include/gemm-config.h) for custom micro-architectural optimization.

## Frequently Asked Questions

### Does BitNet support Apple Silicon?

Yes, Apple Silicon (M1, M2, M3, and later) is fully supported as an ARM64 platform. BitNet automatically enables TL1 kernels with NEON and DOT-PROD extensions when building on Apple Silicon, using 8×256 tiling parameters for optimal matrix multiplication performance. Use the `BITNET_ARM_TL1` CMake option or let [`setup_env.py`](https://github.com/microsoft/BitNet/blob/main/setup_env.py) auto-detect the architecture.

### What is the difference between TL1 and TL2 kernels?

TL2 kernels are optimized specifically for x86-64 processors, utilizing AVX2 and AVX-512 instruction sets with 256-bit and 512-bit vector registers. They employ 4×128 block tiling optimized for Intel and AMD cache hierarchies. TL1 kernels target ARM64 processors using NEON SIMD with optional DOT-PROD extensions, employing larger 8×256 tiling to maximize throughput on ARM memory subsystems.

### Can I run BitNet on older CPUs without AVX2?

BitNet can compile for x86-64 systems with only SSSE3 support, but performance will be significantly degraded compared to AVX2 or AVX-512 paths. The [`setup_env.py`](https://github.com/microsoft/BitNet/blob/main/setup_env.py) script detects available instruction sets and adjusts quantization types accordingly, though the vectorized W2A8 kernels in [`src/ggml-bitnet-mad.cpp`](https://github.com/microsoft/BitNet/blob/main/src/ggml-bitnet-mad.cpp) require at least AVX2 for optimal inference speed.

### How does BitNet detect my CPU architecture during build?

The [`setup_env.py`](https://github.com/microsoft/BitNet/blob/main/setup_env.py) script uses Python's `platform.machine()` function to identify whether the host is `x86_64` or `aarch64` (ARM64). Based on this detection, it injects appropriate compiler flags via the `COMPILER_EXTRA_ARGS` dictionary, selecting either `-DBITNET_X86_TL2=ON` for Intel/AMD or `-DBITNET_ARM_TL1=ON` for ARM processors. The CMakeLists.txt then translates these options into compile definitions that include the correct kernel headers.