Difference Between GEMV and GEMM Operations in BitNet Inference: A Technical Deep Dive
GEMM handles matrix-matrix multiplication for prompt processing with parallel row computation, while GEMV performs matrix-vector multiplication for token generation using single-row activation tensors, both utilizing 2-bit quantized weights in BitNet's inference pipeline.
The microsoft/BitNet repository implements highly optimized inference for 1-bit and 2-bit large language models through two fundamental linear algebra kernels. Understanding the distinction between GEMV and GEMM operations in BitNet inference is essential for optimizing throughput and latency across CPU and GPU deployments.
What Are GEMM and GEMV in BitNet?
GEMM (General Matrix-Matrix Multiply)
GEMM performs multiplication between two dense matrices where both operands contain multiple rows. In BitNet, this operation dominates the prompt processing phase, where the input embedding matrix (containing all prompt tokens) multiplies against the quantized weight matrix.
The operation follows the standard BLAS pattern: C = A × B, where matrix A has dimensions M × K and matrix B has dimensions K × N, producing output C with dimensions M × N. In BitNet's implementation, both M (batch/sequence length) and N (output features) are greater than 1 during prompt encoding.
GEMV (General Matrix-Vector Multiply)
GEMV is a specialized case where the left operand is a single row vector (1 × K) multiplied against a matrix (K × N). BitNet uses GEMV during token generation, where the model processes one new token at a time.
The operation computes c = a × B, where vector a has dimensions 1 × K and matrix B remains K × N. This distinction is critical because the memory access patterns and parallelization strategies differ fundamentally from GEMM, requiring separate optimization paths in the BitNet codebase.
Data Shapes and Computational Patterns
The dimensional constraints determine which kernel BitNet selects during inference:
| Operation | Input A Shape | Input B Shape | Output Shape | Typical Use Case |
|---|---|---|---|---|
| GEMM | M × K (M > 1) | K × N (N > 1) | M × N | Prompt embedding multiplication |
| GEMV | 1 × K | K × N | 1 × N | Single token forward pass |
In src/ggml-bitnet-lut.cpp, the function ggml_bitnet_can_mul_mat validates these dimensional constraints before dispatching to the appropriate kernel implementation.
Implementation Details in the BitNet Codebase
CPU Implementation (ggml-bitnet-mad.cpp)
The core low-precision arithmetic resides in src/ggml-bitnet-mad.cpp, which implements both operations using the same 2-bit weight unpacking primitives but selects different vectorization paths based on the row count parameter nrc (number of rows).
For GEMM operations where multiple rows exist (nrc > 1), the implementation calls ggml_vec_dot_i2_i8_s_1xN, which processes multiple activation rows against the unpacked weight columns in parallel. This path benefits from activation-parallel tiling defined in include/gemm-config.h, where ROW_BLOCK_SIZE, COL_BLOCK_SIZE, and PARALLEL_SIZE control cache blocking and SIMD utilization.
For GEMV operations (nrc = 1), the code selects ggml_vec_dot_i2_i8_s_1x1, found at line 198 of src/ggml-bitnet-mad.cpp. This path iterates over the column dimension without row-level parallelism, using SIMD intrinsics like _mm256_maddubs_epi16 on x86 or NEON equivalents to compute the dot product between the single activation vector and unpacked weight columns.
GPU Implementation (bitnet_kernels.cu)
On NVIDIA GPUs, BitNet provides a dedicated GEMV kernel to maximize token generation throughput. The file gpu/bitnet_kernels/bitnet_kernels.cu contains bitlinear_int8xint2, a specialized CUDA implementation for the M = 1 case.
This kernel launches optimized thread blocks specifically designed for matrix-vector multiplication with 8-bit activations and 2-bit weights, achieving up to 3× speedup over standard BF16 GEMV implementations on A100 GPUs according to the benchmarks in gpu/README.md.
Performance Characteristics and Optimization
GEMM operations dominate the prompt processing phase and benefit from high arithmetic intensity. The tiling strategies in include/gemm-config.h allow the kernel to keep multiple rows of activations and columns of weights in cache simultaneously, maximizing SIMD utilization through ROW_BLOCK_SIZE and COL_BLOCK_SIZE parameters.
GEMV operations are memory-bound rather than compute-bound. Since each token generation step processes only a single activation vector, the kernel cannot amortize weight unpacking costs across multiple rows. The specialized ggml_vec_dot_i2_i8_s_1x1 path and the GPU bitlinear_int8xint2 kernel minimize this overhead by optimizing memory access patterns for single-row traversal.
Summary
- GEMM performs matrix-matrix multiplication for prompt processing, handling multiple input rows (M > 1) with parallel tiling strategies defined in
include/gemm-config.h. - GEMV executes matrix-vector multiplication for token generation, processing single-row activations (M = 1) through specialized paths like
ggml_vec_dot_i2_i8_s_1x1on CPU andbitlinear_int8xint2on GPU. - Both operations utilize the same I2_S 2-bit quantization format but differ in parallelism: GEMM leverages row-parallel tiling while GEMV optimizes for memory-efficient single-row computation.
- The distinction is enforced at runtime through
ggml_bitnet_can_mul_matinsrc/ggml-bitnet-lut.cpp, which dispatches to the appropriate kernel based on input tensor dimensions.
Frequently Asked Questions
When does BitNet use GEMM versus GEMV?
BitNet uses GEMM during the initial prompt encoding phase when processing multiple tokens simultaneously, as the input embedding matrix contains M rows where M equals the prompt length. GEMV activates during the autoregressive token generation phase, where the model processes exactly one new token at a time, resulting in a single-row activation vector. The ggml_bitnet_can_mul_mat function in src/ggml-bitnet-lut.cpp automatically selects the appropriate kernel based on the input tensor's row dimension.
Why is GEMV critical for token generation speed?
GEMV dominates the computational cost during token generation because large language models perform one matrix-vector multiplication per layer for every generated token. Since GEMV operations are memory-bound—limited by the bandwidth to read weights rather than arithmetic capacity—optimizing this kernel directly reduces per-token latency. BitNet's specialized ggml_vec_dot_i2_i8_s_1x1 CPU path and bitlinear_int8xint2 GPU kernel minimize memory overhead for the single-row case, making GEMV optimization essential for real-time inference throughput.
How does the 2-bit quantization affect these operations?
BitNet stores weights in the I2_S format, packing two bits per weight. Both GEMM and GEMV must unpack these weights into 8-bit integers on-the-fly before computing dot products with 8-bit activations. In GEMM, the unpacking cost is amortized across multiple rows of activations processed in parallel, making the overhead negligible relative to computation. In GEMV, the unpacking cost is incurred for every token generation step since only one activation row is processed, making the efficiency of the unpacking primitives in src/ggml-bitnet-mad.cpp critical for overall performance.
Can GEMM and GEMV configurations be tuned for specific hardware?
Yes, GEMM configurations can be tuned through parameters in include/gemm-config.h, specifically ROW_BLOCK_SIZE, COL_BLOCK_SIZE, and PARALLEL_SIZE. These control cache blocking and SIMD parallelism for the matrix-matrix multiplication path, allowing optimization for CPUs with different cache hierarchies or vector widths. GEMV is less configurable because it processes a single row, but the underlying SIMD intrinsics (such as _mm256_maddubs_epi16 on x86 or NEON equivalents) automatically adapt to the target architecture. GPU implementations in gpu/bitnet_kernels/bitnet_kernels.cu can be further optimized by adjusting CUDA block and grid dimensions for specific NVIDIA GPU architectures.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →