Prompt Processing vs Token Generation Benchmarks in BitNet: Understanding GEMM and GEMV Kernels

BitNet separates inference into prompt processing (GEMM matrix-matrix multiplication) and token generation (GEMV matrix-vector multiplication), allowing isolated benchmarking via the -p and -n flags in utils/e2e_benchmark.py.

Microsoft BitNet optimizes large language model inference through two distinct computational kernels that dominate different phases of text generation. Understanding the difference between prompt processing and token generation benchmarks is essential for profiling model performance and identifying bottlenecks in quantized inference. The repository provides specific tooling to measure each phase independently through the llama-bench integration in utils/e2e_benchmark.py.

The Two Phases of BitNet Inference

Prompt Processing with GEMM Operations

During the initial phase, BitNet performs GEMM (General Matrix-Matrix Multiplication) operations to process the entire input prompt at once. The kernel multiplies the input embedding matrix by the model's weight matrices to produce the initial hidden states. According to src/README.md, this phase uses "GEMM Operations: Efficient matrix-matrix multiplication for prompt processing." This computation scales with the number of prompt tokens and executes only once per generation request.

Token Generation with GEMV Operations

Once the prompt is processed, the GEMV (General Matrix-Vector Multiplication) kernel dominates the autoregressive decoding loop. At each generation step, the hidden state vector is multiplied by weight matrices to produce next-token logits. Because only a single token vector is involved, this matrix-vector product executes repeatedly for every generated token. The src/README.md documents this as "GEMV Operations: Optimized matrix-vector multiplication for token generation."

Kernel Implementation in Source Code

The distinct computational patterns are implemented in separate C++ source files within the src/ directory.

BitNet implements prompt processing in src/ggml-bitnet-mad.cpp, which contains the GEMM kernel optimized for processing large input matrices efficiently. For token generation, the repository uses src/ggml-bitnet-lut.cpp, which implements the GEMV kernel used during the autoregressive decoding phase.

These kernels are compiled into the llama-bench binary through the build process defined in utils/tune_gemm_config.py, and are wired into the llama.cpp compute graph to enable accurate performance measurement of each phase.

Isolating Benchmarks with e2e_benchmark.py

The utils/e2e_benchmark.py script provides a CLI interface to the llama-bench binary, exposing the -p (prompt length) and -n (number of tokens to generate) parameters to isolate specific performance characteristics.

To benchmark prompt processing exclusively, set a large prompt size with zero generation tokens:

python utils/e2e_benchmark.py \
    -m models/bitnet-b1_58-large.tl2.gguf \
    -p 2048 \
    -n 0 \
    -t 4

This configuration forces the benchmark to measure only the GEMM operations, as the runtime is completely dominated by the initial matrix-matrix multiplications.

To benchmark token generation performance, minimize the prompt length while maximizing the output length:

python utils/e2e_benchmark.py \
    -m models/bitnet-b1_58-large.tl2.gguf \
    -p 64 \
    -n 512 \
    -t 4

With only 64 prompt tokens, the initial GEMM overhead is negligible compared to the 512 repeated GEMV operations required for autoregressive generation, exposing the per-token latency of the matrix-vector kernel.

For complete end-to-end measurement, balance both parameters:

python utils/e2e_benchmark.py \
    -m models/bitnet-b1_58-large.tl2.gguf \
    -p 512 \
    -n 128 \
    -t 8

The Python script constructs the underlying llama-bench command by forwarding these arguments directly to the benchmark binary:

command = [
    bench_path,
    '-m', args.model,
    '-n', str(args.n_token),
    '-ngl', '0',
    '-b', '1',
    '-t', str(args.threads),
    '-p', str(args.n_prompt),
    '-r', '5'
]

Summary

  • Prompt processing relies on GEMM (matrix-matrix multiplication) implemented in src/ggml-bitnet-mad.cpp, processing all input tokens simultaneously.
  • Token generation utilizes GEMV (matrix-vector multiplication) from src/ggml-bitnet-lut.cpp, executing once per output token during autoregressive decoding.
  • Use utils/e2e_benchmark.py with -n 0 to isolate prompt processing performance, or minimize -p and maximize -n to measure generation throughput.
  • Both kernels are compiled into the llama-bench binary via utils/tune_gemm_config.py, allowing precise profiling of each inference phase.

Frequently Asked Questions

What is the difference between GEMM and GEMV in BitNet benchmarks?

GEMM (General Matrix-Matrix Multiplication) handles prompt processing by multiplying input embedding matrices with weight matrices for the entire prompt at once, while GEMV (General Matrix-Vector Multiplication) handles token generation by performing matrix-vector products for single tokens during autoregressive decoding. The former dominates when processing long inputs, while the latter dominates during extended generation tasks.

How do I benchmark only prompt processing in BitNet?

Run utils/e2e_benchmark.py with a large value for -p (prompt tokens) and -n 0 (zero generated tokens). This configuration forces the benchmark to execute only the initial GEMM operations without entering the autoregressive generation loop, effectively measuring prompt-processing throughput in isolation.

Which source files implement the prompt processing and token generation kernels?

Prompt processing is implemented in src/ggml-bitnet-mad.cpp which contains the GEMM kernel, while token generation is implemented in src/ggml-bitnet-lut.cpp which contains the optimized GEMV kernel. Both files are integrated into the llama.cpp compute graph and compiled into the benchmark binary through the build system.

Why does token generation use matrix-vector multiplication instead of matrix-matrix?

During autoregressive generation, only a single new token is produced at each step, resulting in a vector (not a matrix) that must be multiplied against the weight matrices. This sequential dependency prevents batching the operations as a matrix-matrix product, necessitating the repeated GEMV pattern documented in the BitNet source code.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →