# Prompt Processing vs Token Generation Benchmarks in BitNet: Understanding GEMM and GEMV Kernels

> Explore BitNet prompt processing vs token generation benchmarks. Learn how GEMM and GEMV kernels are used for isolated performance testing with Microsofts BitNet model.

- Repository: [Microsoft/BitNet](https://github.com/microsoft/BitNet)
- Tags: performance
- Published: 2026-03-13

---

**BitNet separates inference into prompt processing (GEMM matrix-matrix multiplication) and token generation (GEMV matrix-vector multiplication), allowing isolated benchmarking via the `-p` and `-n` flags in [`utils/e2e_benchmark.py`](https://github.com/microsoft/BitNet/blob/main/utils/e2e_benchmark.py).**

Microsoft BitNet optimizes large language model inference through two distinct computational kernels that dominate different phases of text generation. Understanding the difference between prompt processing and token generation benchmarks is essential for profiling model performance and identifying bottlenecks in quantized inference. The repository provides specific tooling to measure each phase independently through the `llama-bench` integration in [`utils/e2e_benchmark.py`](https://github.com/microsoft/BitNet/blob/main/utils/e2e_benchmark.py).

## The Two Phases of BitNet Inference

### Prompt Processing with GEMM Operations

During the initial phase, BitNet performs **GEMM** (General Matrix-Matrix Multiplication) operations to process the entire input prompt at once. The kernel multiplies the input embedding matrix by the model's weight matrices to produce the initial hidden states. According to [`src/README.md`](https://github.com/microsoft/BitNet/blob/main/src/README.md), this phase uses "GEMM Operations: Efficient matrix-matrix multiplication for prompt processing." This computation scales with the number of prompt tokens and executes only once per generation request.

### Token Generation with GEMV Operations

Once the prompt is processed, the **GEMV** (General Matrix-Vector Multiplication) kernel dominates the autoregressive decoding loop. At each generation step, the hidden state vector is multiplied by weight matrices to produce next-token logits. Because only a single token vector is involved, this matrix-vector product executes repeatedly for every generated token. The [`src/README.md`](https://github.com/microsoft/BitNet/blob/main/src/README.md) documents this as "GEMV Operations: Optimized matrix-vector multiplication for token generation."

## Kernel Implementation in Source Code

The distinct computational patterns are implemented in separate C++ source files within the `src/` directory.

BitNet implements prompt processing in [`src/ggml-bitnet-mad.cpp`](https://github.com/microsoft/BitNet/blob/main/src/ggml-bitnet-mad.cpp), which contains the GEMM kernel optimized for processing large input matrices efficiently. For token generation, the repository uses [`src/ggml-bitnet-lut.cpp`](https://github.com/microsoft/BitNet/blob/main/src/ggml-bitnet-lut.cpp), which implements the GEMV kernel used during the autoregressive decoding phase.

These kernels are compiled into the `llama-bench` binary through the build process defined in [`utils/tune_gemm_config.py`](https://github.com/microsoft/BitNet/blob/main/utils/tune_gemm_config.py), and are wired into the llama.cpp compute graph to enable accurate performance measurement of each phase.

## Isolating Benchmarks with e2e_benchmark.py

The [`utils/e2e_benchmark.py`](https://github.com/microsoft/BitNet/blob/main/utils/e2e_benchmark.py) script provides a CLI interface to the `llama-bench` binary, exposing the `-p` (prompt length) and `-n` (number of tokens to generate) parameters to isolate specific performance characteristics.

To benchmark prompt processing exclusively, set a large prompt size with zero generation tokens:

```bash
python utils/e2e_benchmark.py \
    -m models/bitnet-b1_58-large.tl2.gguf \
    -p 2048 \
    -n 0 \
    -t 4

```

This configuration forces the benchmark to measure only the GEMM operations, as the runtime is completely dominated by the initial matrix-matrix multiplications.

To benchmark token generation performance, minimize the prompt length while maximizing the output length:

```bash
python utils/e2e_benchmark.py \
    -m models/bitnet-b1_58-large.tl2.gguf \
    -p 64 \
    -n 512 \
    -t 4

```

With only 64 prompt tokens, the initial GEMM overhead is negligible compared to the 512 repeated GEMV operations required for autoregressive generation, exposing the per-token latency of the matrix-vector kernel.

For complete end-to-end measurement, balance both parameters:

```bash
python utils/e2e_benchmark.py \
    -m models/bitnet-b1_58-large.tl2.gguf \
    -p 512 \
    -n 128 \
    -t 8

```

The Python script constructs the underlying `llama-bench` command by forwarding these arguments directly to the benchmark binary:

```python
command = [
    bench_path,
    '-m', args.model,
    '-n', str(args.n_token),
    '-ngl', '0',
    '-b', '1',
    '-t', str(args.threads),
    '-p', str(args.n_prompt),
    '-r', '5'
]

```

## Summary

- **Prompt processing** relies on GEMM (matrix-matrix multiplication) implemented in [`src/ggml-bitnet-mad.cpp`](https://github.com/microsoft/BitNet/blob/main/src/ggml-bitnet-mad.cpp), processing all input tokens simultaneously.
- **Token generation** utilizes GEMV (matrix-vector multiplication) from [`src/ggml-bitnet-lut.cpp`](https://github.com/microsoft/BitNet/blob/main/src/ggml-bitnet-lut.cpp), executing once per output token during autoregressive decoding.
- Use [`utils/e2e_benchmark.py`](https://github.com/microsoft/BitNet/blob/main/utils/e2e_benchmark.py) with `-n 0` to isolate prompt processing performance, or minimize `-p` and maximize `-n` to measure generation throughput.
- Both kernels are compiled into the `llama-bench` binary via [`utils/tune_gemm_config.py`](https://github.com/microsoft/BitNet/blob/main/utils/tune_gemm_config.py), allowing precise profiling of each inference phase.

## Frequently Asked Questions

### What is the difference between GEMM and GEMV in BitNet benchmarks?

GEMM (General Matrix-Matrix Multiplication) handles prompt processing by multiplying input embedding matrices with weight matrices for the entire prompt at once, while GEMV (General Matrix-Vector Multiplication) handles token generation by performing matrix-vector products for single tokens during autoregressive decoding. The former dominates when processing long inputs, while the latter dominates during extended generation tasks.

### How do I benchmark only prompt processing in BitNet?

Run [`utils/e2e_benchmark.py`](https://github.com/microsoft/BitNet/blob/main/utils/e2e_benchmark.py) with a large value for `-p` (prompt tokens) and `-n 0` (zero generated tokens). This configuration forces the benchmark to execute only the initial GEMM operations without entering the autoregressive generation loop, effectively measuring prompt-processing throughput in isolation.

### Which source files implement the prompt processing and token generation kernels?

Prompt processing is implemented in [`src/ggml-bitnet-mad.cpp`](https://github.com/microsoft/BitNet/blob/main/src/ggml-bitnet-mad.cpp) which contains the GEMM kernel, while token generation is implemented in [`src/ggml-bitnet-lut.cpp`](https://github.com/microsoft/BitNet/blob/main/src/ggml-bitnet-lut.cpp) which contains the optimized GEMV kernel. Both files are integrated into the llama.cpp compute graph and compiled into the benchmark binary through the build system.

### Why does token generation use matrix-vector multiplication instead of matrix-matrix?

During autoregressive generation, only a single new token is produced at each step, resulting in a vector (not a matrix) that must be multiplied against the weight matrices. This sequential dependency prevents batching the operations as a matrix-matrix product, necessitating the repeated GEMV pattern documented in the BitNet source code.