# Weight Parallel vs Activation Parallel in BitNet Kernels: A Complete Guide

> Understand weight parallel vs activation parallel in BitNet kernels. Learn how BitNet optimizes performance by processing weights and activations efficiently for faster computation.

- Repository: [Microsoft/BitNet](https://github.com/microsoft/BitNet)
- Tags: deep-dive
- Published: 2026-03-13

---

**Weight parallel processes multiple weight rows per kernel launch to reduce overhead, while activation parallel amortizes the cost of unpacking 2-bit I2_S weights by computing multiple activation elements simultaneously.**

In the `microsoft/BitNet` repository, these two parallelization strategies determine how quantized matrix multiplication kernels execute on CPU. Understanding **weight parallel vs activation parallel** is essential for optimizing inference performance, particularly when working with BitNet's 1.58-bit or 2-bit quantized models.

## What Are Weight Parallel and Activation Parallel?

BitNet's GEMM (General Matrix Multiply) kernels support two distinct parallelization modes controlled by the `ACT_PARALLEL` macro in [`include/gemm-config.h`](https://github.com/microsoft/BitNet/blob/main/include/gemm-config.h). These modes dictate how the kernel distributes work across CPU threads and how it handles the computational overhead of unpacking quantized weights.

- **Weight Parallel**: Optimizes for scenarios where weight data requires minimal preprocessing, grouping multiple weight elements per kernel launch to amortize invocation overhead.
- **Activation Parallel**: Optimizes for the I2_S quantization format by interleaving weight unpacking with activation computation, hiding latency by processing multiple activation elements concurrently.

## Weight Parallel Mode

### How Weight Parallel Works

In weight parallel mode, the kernel processes blocks of the weight matrix in a single invocation. According to the configuration in [`include/gemm-config.h`](https://github.com/microsoft/BitNet/blob/main/include/gemm-config.h), when `ACT_PARALLEL` is undefined, the kernel uses these block dimensions:

```c
#define ROW_BLOCK_SIZE 128
#define COL_BLOCK_SIZE 32
#define PARALLEL_SIZE 8

```

The implementation in [`src/ggml-bitnet-mad.cpp`](https://github.com/microsoft/BitNet/blob/main/src/ggml-bitnet-mad.cpp) iterates over these weight blocks, computing dot products with a single activation vector per launch. By handling 128 rows and 32 columns of weights simultaneously, the kernel minimizes the fixed cost of kernel invocation.

### When to Use Weight Parallel

Weight parallel excels when working with formats that have low unpacking overhead:

- **FP16 or Q8 quantized weights**: These formats store weights in a "ready-to-use" state where the cost of unpacking is negligible.
- **High-throughput scenarios**: When memory bandwidth is not constrained by dequantization operations, the reduced kernel launch overhead provides measurable gains.

## Activation Parallel Mode

### How Activation Parallel Works

Activation parallel builds upon weight parallel but adds a critical optimization for packed quantization formats. When `ACT_PARALLEL` is defined in [`include/gemm-config.h`](https://github.com/microsoft/BitNet/blob/main/include/gemm-config.h), the block dimensions shift to:

```c
#define ROW_BLOCK_SIZE 4
#define COL_BLOCK_SIZE 128
#define PARALLEL_SIZE 4

```

The I2_S format stores weights in a packed 2-bit representation that must be unpacked before use. In activation parallel mode, the kernel in [`src/ggml-bitnet-mad.cpp`](https://github.com/microsoft/BitNet/blob/main/src/ggml-bitnet-mad.cpp) spreads this unpacking work across multiple activation elements. While unpacking a weight block, the kernel simultaneously computes results for four activation rows and 128 columns, keeping the CPU's execution units busy during the unpacking latency.

### Why Activation Parallel is the Default

The `microsoft/BitNet` repository enables activation parallel by default because:

1. **I2_S dominates runtime**: For 2-bit quantized models, the unpacking step becomes the primary bottleneck. Activation parallel converts this overhead into useful computation.
2. **Measurable speedup**: The performance tables in [`src/README.md`](https://github.com/microsoft/BitNet/blob/main/src/README.md) demonstrate that activation parallel provides significant throughput improvements for I2_S inference compared to weight parallel.

## Configuring the Parallel Mode in BitNet

You control the parallelization strategy through the `ACT_PARALLEL` macro in [`include/gemm-config.h`](https://github.com/microsoft/BitNet/blob/main/include/gemm-config.h).

To verify or change the current mode, examine the header file:

```c
/* include/gemm-config.h */
#define ACT_PARALLEL          // Comment this out to use weight parallel

#if defined(__AVX__) || defined(__AVX2__) || defined(__AVX512F__) || defined(__SSSE3__)
#if defined(ACT_PARALLEL)
    #define ROW_BLOCK_SIZE 4
    #define COL_BLOCK_SIZE 128
    #define PARALLEL_SIZE 4
#else
    #define ROW_BLOCK_SIZE 128
    #define COL_BLOCK_SIZE 32
    #define PARALLEL_SIZE 8
#endif
#endif

```

To build with weight parallel instead of the default activation parallel:

```bash

# Edit include/gemm-config.h to comment out ACT_PARALLEL

# Then rebuild

cmake -S . -B build -DCMAKE_BUILD_TYPE=Release
make -C build

```

The kernel implementation in [`src/ggml-bitnet-mad.cpp`](https://github.com/microsoft/BitNet/blob/main/src/ggml-bitnet-mad.cpp) uses conditional compilation to select the appropriate code path based on this macro.

## Performance Implications

Choosing between weight parallel and activation parallel involves trade-offs between kernel launch overhead and dequantization latency:

- **Weight parallel** minimizes CPU dispatch overhead by processing large weight blocks (128×32) per invocation. This benefits formats like FP16 where memory bandwidth, not computation, limits performance.
- **Activation parallel** accepts smaller weight blocks (4×128) but hides the I2_S unpacking cost behind parallel activation computation. This transforms dequantization from a bottleneck into overlapped work.

For I2_S quantized models in `microsoft/BitNet`, activation parallel typically delivers superior throughput because the 2-bit unpacking operation would otherwise stall the pipeline.

## Summary

- **Weight parallel** processes multiple weight rows per kernel launch (128×32 blocks), reducing invocation overhead for formats with minimal unpacking costs like FP16 or Q8.
- **Activation parallel** processes smaller weight blocks (4×128) while computing multiple activation elements simultaneously, hiding the latency of unpacking I2_S 2-bit weights.
- The mode is controlled by the `ACT_PARALLEL` macro in [`include/gemm-config.h`](https://github.com/microsoft/BitNet/blob/main/include/gemm-config.h), with activation parallel enabled by default for optimal I2_S performance.
- Implementation details reside in [`src/ggml-bitnet-mad.cpp`](https://github.com/microsoft/BitNet/blob/main/src/ggml-bitnet-mad.cpp), where conditional compilation selects the appropriate GEMM strategy based on the configuration header.

## Frequently Asked Questions

### Which parallel mode should I use for I2_S quantization?

You should use **activation parallel** for I2_S quantization. This mode is the default in [`include/gemm-config.h`](https://github.com/microsoft/BitNet/blob/main/include/gemm-config.h) because the I2_S format stores weights in packed 2-bit representations that require expensive unpacking. Activation parallel amortizes this cost by processing multiple activation elements (4 rows × 128 columns) while the unpacking occurs, delivering measurable speedups over weight parallel for 2-bit models.

### How do I switch from activation parallel to weight parallel?

To switch modes, edit [`include/gemm-config.h`](https://github.com/microsoft/BitNet/blob/main/include/gemm-config.h) and comment out or remove the line `#define ACT_PARALLEL`. This forces the build system to use the weight parallel block sizes (128×32) instead of the activation parallel dimensions (4×128). After modifying the header, rebuild the project with `cmake` and `make` to compile the weight parallel implementation in [`src/ggml-bitnet-mad.cpp`](https://github.com/microsoft/BitNet/blob/main/src/ggml-bitnet-mad.cpp).

### What are the block size differences between weight parallel and activation parallel?

Weight parallel uses larger blocks of **128 rows × 32 columns** with a parallel size of 8, optimized for processing large weight chunks per kernel invocation. Activation parallel uses smaller blocks of **4 rows × 128 columns** with a parallel size of 4, designed to keep CPU execution units busy during the I2_S weight unpacking process. These constants are defined in [`include/gemm-config.h`](https://github.com/microsoft/BitNet/blob/main/include/gemm-config.h) and consumed by the GEMM kernels in [`src/ggml-bitnet-mad.cpp`](https://github.com/microsoft/BitNet/blob/main/src/ggml-bitnet-mad.cpp).

### Why does activation parallel help with 2-bit weights?

Activation parallel helps because **2-bit I2_S weights require unpacking** before they can be used in dot-product calculations. This unpacking introduces memory traffic and latency that would otherwise stall the CPU pipeline. By processing four activation rows simultaneously while unpacking the weight data, activation parallel converts this unpacking overhead into useful computation, effectively hiding the latency and improving throughput for 2-bit quantized models.