Weight Parallel vs Activation Parallel in BitNet Kernels: A Complete Guide

Weight parallel processes multiple weight rows per kernel launch to reduce overhead, while activation parallel amortizes the cost of unpacking 2-bit I2_S weights by computing multiple activation elements simultaneously.

In the microsoft/BitNet repository, these two parallelization strategies determine how quantized matrix multiplication kernels execute on CPU. Understanding weight parallel vs activation parallel is essential for optimizing inference performance, particularly when working with BitNet's 1.58-bit or 2-bit quantized models.

What Are Weight Parallel and Activation Parallel?

BitNet's GEMM (General Matrix Multiply) kernels support two distinct parallelization modes controlled by the ACT_PARALLEL macro in include/gemm-config.h. These modes dictate how the kernel distributes work across CPU threads and how it handles the computational overhead of unpacking quantized weights.

  • Weight Parallel: Optimizes for scenarios where weight data requires minimal preprocessing, grouping multiple weight elements per kernel launch to amortize invocation overhead.
  • Activation Parallel: Optimizes for the I2_S quantization format by interleaving weight unpacking with activation computation, hiding latency by processing multiple activation elements concurrently.

Weight Parallel Mode

How Weight Parallel Works

In weight parallel mode, the kernel processes blocks of the weight matrix in a single invocation. According to the configuration in include/gemm-config.h, when ACT_PARALLEL is undefined, the kernel uses these block dimensions:

#define ROW_BLOCK_SIZE 128
#define COL_BLOCK_SIZE 32
#define PARALLEL_SIZE 8

The implementation in src/ggml-bitnet-mad.cpp iterates over these weight blocks, computing dot products with a single activation vector per launch. By handling 128 rows and 32 columns of weights simultaneously, the kernel minimizes the fixed cost of kernel invocation.

When to Use Weight Parallel

Weight parallel excels when working with formats that have low unpacking overhead:

  • FP16 or Q8 quantized weights: These formats store weights in a "ready-to-use" state where the cost of unpacking is negligible.
  • High-throughput scenarios: When memory bandwidth is not constrained by dequantization operations, the reduced kernel launch overhead provides measurable gains.

Activation Parallel Mode

How Activation Parallel Works

Activation parallel builds upon weight parallel but adds a critical optimization for packed quantization formats. When ACT_PARALLEL is defined in include/gemm-config.h, the block dimensions shift to:

#define ROW_BLOCK_SIZE 4
#define COL_BLOCK_SIZE 128
#define PARALLEL_SIZE 4

The I2_S format stores weights in a packed 2-bit representation that must be unpacked before use. In activation parallel mode, the kernel in src/ggml-bitnet-mad.cpp spreads this unpacking work across multiple activation elements. While unpacking a weight block, the kernel simultaneously computes results for four activation rows and 128 columns, keeping the CPU's execution units busy during the unpacking latency.

Why Activation Parallel is the Default

The microsoft/BitNet repository enables activation parallel by default because:

  1. I2_S dominates runtime: For 2-bit quantized models, the unpacking step becomes the primary bottleneck. Activation parallel converts this overhead into useful computation.
  2. Measurable speedup: The performance tables in src/README.md demonstrate that activation parallel provides significant throughput improvements for I2_S inference compared to weight parallel.

Configuring the Parallel Mode in BitNet

You control the parallelization strategy through the ACT_PARALLEL macro in include/gemm-config.h.

To verify or change the current mode, examine the header file:

/* include/gemm-config.h */
#define ACT_PARALLEL          // Comment this out to use weight parallel

#if defined(__AVX__) || defined(__AVX2__) || defined(__AVX512F__) || defined(__SSSE3__)
#if defined(ACT_PARALLEL)
    #define ROW_BLOCK_SIZE 4
    #define COL_BLOCK_SIZE 128
    #define PARALLEL_SIZE 4
#else
    #define ROW_BLOCK_SIZE 128
    #define COL_BLOCK_SIZE 32
    #define PARALLEL_SIZE 8
#endif
#endif

To build with weight parallel instead of the default activation parallel:


# Edit include/gemm-config.h to comment out ACT_PARALLEL

# Then rebuild

cmake -S . -B build -DCMAKE_BUILD_TYPE=Release
make -C build

The kernel implementation in src/ggml-bitnet-mad.cpp uses conditional compilation to select the appropriate code path based on this macro.

Performance Implications

Choosing between weight parallel and activation parallel involves trade-offs between kernel launch overhead and dequantization latency:

  • Weight parallel minimizes CPU dispatch overhead by processing large weight blocks (128×32) per invocation. This benefits formats like FP16 where memory bandwidth, not computation, limits performance.
  • Activation parallel accepts smaller weight blocks (4×128) but hides the I2_S unpacking cost behind parallel activation computation. This transforms dequantization from a bottleneck into overlapped work.

For I2_S quantized models in microsoft/BitNet, activation parallel typically delivers superior throughput because the 2-bit unpacking operation would otherwise stall the pipeline.

Summary

  • Weight parallel processes multiple weight rows per kernel launch (128×32 blocks), reducing invocation overhead for formats with minimal unpacking costs like FP16 or Q8.
  • Activation parallel processes smaller weight blocks (4×128) while computing multiple activation elements simultaneously, hiding the latency of unpacking I2_S 2-bit weights.
  • The mode is controlled by the ACT_PARALLEL macro in include/gemm-config.h, with activation parallel enabled by default for optimal I2_S performance.
  • Implementation details reside in src/ggml-bitnet-mad.cpp, where conditional compilation selects the appropriate GEMM strategy based on the configuration header.

Frequently Asked Questions

Which parallel mode should I use for I2_S quantization?

You should use activation parallel for I2_S quantization. This mode is the default in include/gemm-config.h because the I2_S format stores weights in packed 2-bit representations that require expensive unpacking. Activation parallel amortizes this cost by processing multiple activation elements (4 rows × 128 columns) while the unpacking occurs, delivering measurable speedups over weight parallel for 2-bit models.

How do I switch from activation parallel to weight parallel?

To switch modes, edit include/gemm-config.h and comment out or remove the line #define ACT_PARALLEL. This forces the build system to use the weight parallel block sizes (128×32) instead of the activation parallel dimensions (4×128). After modifying the header, rebuild the project with cmake and make to compile the weight parallel implementation in src/ggml-bitnet-mad.cpp.

What are the block size differences between weight parallel and activation parallel?

Weight parallel uses larger blocks of 128 rows × 32 columns with a parallel size of 8, optimized for processing large weight chunks per kernel invocation. Activation parallel uses smaller blocks of 4 rows × 128 columns with a parallel size of 4, designed to keep CPU execution units busy during the I2_S weight unpacking process. These constants are defined in include/gemm-config.h and consumed by the GEMM kernels in src/ggml-bitnet-mad.cpp.

Why does activation parallel help with 2-bit weights?

Activation parallel helps because 2-bit I2_S weights require unpacking before they can be used in dot-product calculations. This unpacking introduces memory traffic and latency that would otherwise stall the CPU pipeline. By processing four activation rows simultaneously while unpacking the weight data, activation parallel converts this unpacking overhead into useful computation, effectively hiding the latency and improving throughput for 2-bit quantized models.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →