# Difference Between GEMV and GEMM Operations in BitNet Inference: A Technical Deep Dive

> Understand GEMV vs GEMM in BitNet inference. GEMM handles prompt processing matrix-matrix multiplication, while GEMV manages token generation matrix-vector multiplication with 2-bit quantization.

- Repository: [Microsoft/BitNet](https://github.com/microsoft/BitNet)
- Tags: deep-dive
- Published: 2026-03-13

---

**GEMM handles matrix-matrix multiplication for prompt processing with parallel row computation, while GEMV performs matrix-vector multiplication for token generation using single-row activation tensors, both utilizing 2-bit quantized weights in BitNet's inference pipeline.**

The `microsoft/BitNet` repository implements highly optimized inference for 1-bit and 2-bit large language models through two fundamental linear algebra kernels. Understanding the distinction between GEMV and GEMM operations in BitNet inference is essential for optimizing throughput and latency across CPU and GPU deployments.

## What Are GEMM and GEMV in BitNet?

### GEMM (General Matrix-Matrix Multiply)

GEMM performs multiplication between two dense matrices where both operands contain multiple rows. In BitNet, this operation dominates the **prompt processing phase**, where the input embedding matrix (containing all prompt tokens) multiplies against the quantized weight matrix.

The operation follows the standard BLAS pattern: **C = A × B**, where matrix **A** has dimensions *M × K* and matrix **B** has dimensions *K × N*, producing output **C** with dimensions *M × N*. In BitNet's implementation, both **M** (batch/sequence length) and **N** (output features) are greater than 1 during prompt encoding.

### GEMV (General Matrix-Vector Multiply)

GEMV is a specialized case where the left operand is a single row vector (1 × K) multiplied against a matrix (K × N). BitNet uses GEMV during **token generation**, where the model processes one new token at a time.

The operation computes **c = a × B**, where vector **a** has dimensions *1 × K* and matrix **B** remains *K × N*. This distinction is critical because the memory access patterns and parallelization strategies differ fundamentally from GEMM, requiring separate optimization paths in the BitNet codebase.

## Data Shapes and Computational Patterns

The dimensional constraints determine which kernel BitNet selects during inference:

| Operation | Input A Shape | Input B Shape | Output Shape | Typical Use Case |
|-----------|---------------|---------------|--------------|----------------|
| **GEMM** | *M × K* (M > 1) | *K × N* (N > 1) | *M × N* | Prompt embedding multiplication |
| **GEMV** | *1 × K* | *K × N* | *1 × N* | Single token forward pass |

In [`src/ggml-bitnet-lut.cpp`](https://github.com/microsoft/BitNet/blob/main/src/ggml-bitnet-lut.cpp), the function `ggml_bitnet_can_mul_mat` validates these dimensional constraints before dispatching to the appropriate kernel implementation.

## Implementation Details in the BitNet Codebase

### CPU Implementation (ggml-bitnet-mad.cpp)

The core low-precision arithmetic resides in [`src/ggml-bitnet-mad.cpp`](https://github.com/microsoft/BitNet/blob/main/src/ggml-bitnet-mad.cpp), which implements both operations using the same 2-bit weight unpacking primitives but selects different vectorization paths based on the row count parameter `nrc` (number of rows).

For **GEMM** operations where multiple rows exist (`nrc > 1`), the implementation calls `ggml_vec_dot_i2_i8_s_1xN`, which processes multiple activation rows against the unpacked weight columns in parallel. This path benefits from activation-parallel tiling defined in [`include/gemm-config.h`](https://github.com/microsoft/BitNet/blob/main/include/gemm-config.h), where `ROW_BLOCK_SIZE`, `COL_BLOCK_SIZE`, and `PARALLEL_SIZE` control cache blocking and SIMD utilization.

For **GEMV** operations (`nrc = 1`), the code selects `ggml_vec_dot_i2_i8_s_1x1`, found at line 198 of [`src/ggml-bitnet-mad.cpp`](https://github.com/microsoft/BitNet/blob/main/src/ggml-bitnet-mad.cpp). This path iterates over the column dimension without row-level parallelism, using SIMD intrinsics like `_mm256_maddubs_epi16` on x86 or NEON equivalents to compute the dot product between the single activation vector and unpacked weight columns.

### GPU Implementation (bitnet_kernels.cu)

On NVIDIA GPUs, BitNet provides a dedicated GEMV kernel to maximize token generation throughput. The file `gpu/bitnet_kernels/bitnet_kernels.cu` contains `bitlinear_int8xint2`, a specialized CUDA implementation for the *M = 1* case.

This kernel launches optimized thread blocks specifically designed for matrix-vector multiplication with 8-bit activations and 2-bit weights, achieving up to 3× speedup over standard BF16 GEMV implementations on A100 GPUs according to the benchmarks in [`gpu/README.md`](https://github.com/microsoft/BitNet/blob/main/gpu/README.md).

## Performance Characteristics and Optimization

**GEMM** operations dominate the prompt processing phase and benefit from high arithmetic intensity. The tiling strategies in [`include/gemm-config.h`](https://github.com/microsoft/BitNet/blob/main/include/gemm-config.h) allow the kernel to keep multiple rows of activations and columns of weights in cache simultaneously, maximizing SIMD utilization through `ROW_BLOCK_SIZE` and `COL_BLOCK_SIZE` parameters.

**GEMV** operations are memory-bound rather than compute-bound. Since each token generation step processes only a single activation vector, the kernel cannot amortize weight unpacking costs across multiple rows. The specialized `ggml_vec_dot_i2_i8_s_1x1` path and the GPU `bitlinear_int8xint2` kernel minimize this overhead by optimizing memory access patterns for single-row traversal.

## Summary

- **GEMM** performs matrix-matrix multiplication for prompt processing, handling multiple input rows (*M > 1*) with parallel tiling strategies defined in [`include/gemm-config.h`](https://github.com/microsoft/BitNet/blob/main/include/gemm-config.h).
- **GEMV** executes matrix-vector multiplication for token generation, processing single-row activations (*M = 1*) through specialized paths like `ggml_vec_dot_i2_i8_s_1x1` on CPU and `bitlinear_int8xint2` on GPU.
- Both operations utilize the same **I2_S** 2-bit quantization format but differ in parallelism: GEMM leverages row-parallel tiling while GEMV optimizes for memory-efficient single-row computation.
- The distinction is enforced at runtime through `ggml_bitnet_can_mul_mat` in [`src/ggml-bitnet-lut.cpp`](https://github.com/microsoft/BitNet/blob/main/src/ggml-bitnet-lut.cpp), which dispatches to the appropriate kernel based on input tensor dimensions.

## Frequently Asked Questions

### When does BitNet use GEMM versus GEMV?

BitNet uses **GEMM** during the initial prompt encoding phase when processing multiple tokens simultaneously, as the input embedding matrix contains *M* rows where *M* equals the prompt length. **GEMV** activates during the autoregressive token generation phase, where the model processes exactly one new token at a time, resulting in a single-row activation vector. The `ggml_bitnet_can_mul_mat` function in [`src/ggml-bitnet-lut.cpp`](https://github.com/microsoft/BitNet/blob/main/src/ggml-bitnet-lut.cpp) automatically selects the appropriate kernel based on the input tensor's row dimension.

### Why is GEMV critical for token generation speed?

GEMV dominates the computational cost during token generation because large language models perform one matrix-vector multiplication per layer for every generated token. Since GEMV operations are memory-bound—limited by the bandwidth to read weights rather than arithmetic capacity—optimizing this kernel directly reduces per-token latency. BitNet's specialized `ggml_vec_dot_i2_i8_s_1x1` CPU path and `bitlinear_int8xint2` GPU kernel minimize memory overhead for the single-row case, making GEMV optimization essential for real-time inference throughput.

### How does the 2-bit quantization affect these operations?

BitNet stores weights in the **I2_S** format, packing two bits per weight. Both GEMM and GEMV must unpack these weights into 8-bit integers on-the-fly before computing dot products with 8-bit activations. In GEMM, the unpacking cost is amortized across multiple rows of activations processed in parallel, making the overhead negligible relative to computation. In GEMV, the unpacking cost is incurred for every token generation step since only one activation row is processed, making the efficiency of the unpacking primitives in [`src/ggml-bitnet-mad.cpp`](https://github.com/microsoft/BitNet/blob/main/src/ggml-bitnet-mad.cpp) critical for overall performance.

### Can GEMM and GEMV configurations be tuned for specific hardware?

Yes, GEMM configurations can be tuned through parameters in [`include/gemm-config.h`](https://github.com/microsoft/BitNet/blob/main/include/gemm-config.h), specifically `ROW_BLOCK_SIZE`, `COL_BLOCK_SIZE`, and `PARALLEL_SIZE`. These control cache blocking and SIMD parallelism for the matrix-matrix multiplication path, allowing optimization for CPUs with different cache hierarchies or vector widths. GEMV is less configurable because it processes a single row, but the underlying SIMD intrinsics (such as `_mm256_maddubs_epi16` on x86 or NEON equivalents) automatically adapt to the target architecture. GPU implementations in `gpu/bitnet_kernels/bitnet_kernels.cu` can be further optimized by adjusting CUDA block and grid dimensions for specific NVIDIA GPU architectures.