# How to Benchmark BitNet Inference Performance

> Benchmark BitNet inference performance using llama-bench. Measure tokens-per-second across various thread counts, prompt lengths, and model configurations with our easy-to-use script.

- Repository: [Microsoft/BitNet](https://github.com/microsoft/BitNet)
- Tags: performance
- Published: 2026-03-13

---

**Use the [`utils/e2e_benchmark.py`](https://github.com/microsoft/BitNet/blob/main/utils/e2e_benchmark.py) script to drive the compiled `llama-bench` binary, measuring tokens-per-second (tps) across different thread counts, prompt lengths, and model configurations.**

Microsoft/BitNet ships with a built-in end-to-end benchmarking system designed to standardize performance measurement across x86 and ARM platforms. To accurately benchmark BitNet inference performance, you will compile the optimized C++ kernels, prepare a quantized GGUF model, and execute the Python wrapper that orchestrates the `llama-bench` binary with consistent parameters.

## Architecture of the Benchmark System

The benchmarking workflow relies on a Python frontend and a high-performance C++ backend. Understanding how these components interact ensures you can troubleshoot issues and customize measurements.

### Binary Discovery and Cross-Platform Handling

In [`utils/e2e_benchmark.py`](https://github.com/microsoft/BitNet/blob/main/utils/e2e_benchmark.py), lines **L25-L33** implement automatic discovery of the benchmark executable. The script searches the `build/bin` directory and automatically selects the correct binary name for your platform—`llama-bench.exe` on Windows or `llama-bench` on Unix-like systems—eliminating manual path configuration.

### Command Construction and Execution

Lines **L36-L45** of [`utils/e2e_benchmark.py`](https://github.com/microsoft/BitNet/blob/main/utils/e2e_benchmark.py) assemble the command-line arguments for the benchmark binary. The script injects the model path (`-m`), generation length (`-n`), prompt size (`-p`), thread count (`-t`), and a fixed repeat count of five (`-r 5`) to ensure statistical significance. Lines **L8-L18** handle execution via `subprocess.run`, capturing both stdout and stderr to a timestamped log file in your specified `log_dir` (defaulting to the current working directory).

### Core Inference Kernels

The speed advantages of BitNet stem from specialized 1-bit matrix operations implemented in the C++ source. The **Multiply-Add (MAD)** kernel lives in [`src/ggml-bitnet-mad.cpp`](https://github.com/microsoft/BitNet/blob/main/src/ggml-bitnet-mad.cpp) (lines **L1-L20**), while the **Lookup-Table (LUT)** kernel resides in [`src/ggml-bitnet-lut.cpp`](https://github.com/microsoft/BitNet/blob/main/src/ggml-bitnet-lut.cpp) (lines **L1-L20**). These kernels replace traditional floating-point multiplication with bitwise operations, enabling the high throughput measured by the benchmark.

## Step-by-Step Benchmarking Workflow

Follow these steps to generate reproducible performance metrics for any BitNet model.

1. **Build the project in release mode**

   Compile the `llama-bench` binary and core libraries with optimizations enabled. Release mode is critical for realistic throughput numbers.

   ```bash
   git clone --recursive https://github.com/microsoft/BitNet.git
   cd BitNet
   mkdir -p build && cd build
   cmake -DCMAKE_BUILD_TYPE=Release .. && make -j$(nproc)
   ```

2. **Download and prepare a BitNet model**

   Retrieve a quantized GGUF model from Hugging Face and run the setup script to prepare the environment.

   ```bash
   huggingface-cli download microsoft/BitNet-b1.58-2B-4T --local-dir models/BitNet-2B
   python setup_env.py -md models/BitNet-2B -q i2_s
   ```

3. **Run the end-to-end benchmark**

   Execute [`utils/e2e_benchmark.py`](https://github.com/microsoft/BitNet/blob/main/utils/e2e_benchmark.py) with your model path and desired test parameters. The following example generates 200 tokens from a 256-token prompt using 8 threads:

   ```bash
   python utils/e2e_benchmark.py \
       -m models/BitNet-2B/ggml-model-i2_s.gguf \
       -n 200 \
       -p 256 \
       -t 8
   ```

   The script repeats the test five times and writes raw `llama-bench` output to a log file named `e2e_benchmark_<timestamp>.log`.

4. **Inspect the results**

   Read the generated log to extract the tokens-per-second (tps) metric. Look for lines containing `tps:` followed by the numerical value.

   ```bash
   cat e2e_benchmark_$(date +%s).log
   ```

   Typical output includes per-iteration timing and an aggregate tps average across the five repeats.

5. **(Optional) Measure single-prompt latency**

   For wall-clock latency measurements of individual requests, use the high-level Python wrapper instead of the throughput benchmark:

   ```bash
   time python run_inference.py \
       -m models/BitNet-2B/ggml-model-i2_s.gguf \
       -p "Explain quantum computing" \
       -t 8
   ```

## Interpreting Benchmark Results

**Tokens-per-second (tps)** is the primary metric for throughput. Higher values indicate faster inference. When you benchmark BitNet inference performance, compare results across these dimensions:

- **Quantization schemes**: Test `i2_s` (2-bit), `tl1`, and `tl2` (ternary 1.58-bit) configurations to evaluate the speed-precision trade-off. Lower bit widths generally yield higher tps on CPU.
- **Thread scaling**: Increase the `-t` parameter to measure how well the MAD and LUT kernels utilize multi-core x86 or ARM processors. Reference the performance tables in `src/assets/performance.png` (cited in [`README.md`](https://github.com/microsoft/BitNet/blob/main/README.md) lines **L15-L18**) for expected scaling curves on specific hardware.
- **Model size impact**: Compare the 2B parameter model against larger variants to quantify the overhead of increased model capacity versus the efficiency of 1-bit compute.

## Best Practices for Accurate Results

To avoid skewed measurements when using [`utils/e2e_benchmark.py`](https://github.com/microsoft/BitNet/blob/main/utils/e2e_benchmark.py), adhere to these guidelines:

- **Always build in release mode**. Debug builds (`-DCMAKE_BUILD_TYPE=Debug`) disable compiler optimizations and yield unrealistically low tps.
- **Disable CPU frequency scaling**. Power-saving features like Intel SpeedStep or AMD Cool'n'Quiet can throttle cores during the benchmark run, creating variance between iterations.
- **Account for warm-up effects**. The script executes five repeats (`-r 5`); discard the first iteration to allow CPU caches and branch predictors to stabilize, then average the remaining four runs.
- **Isolate the test environment**. Close background applications that compete for CPU cycles or memory bandwidth to ensure the `llama-bench` binary has exclusive resource access.

## Summary

- Use [`utils/e2e_benchmark.py`](https://github.com/microsoft/BitNet/blob/main/utils/e2e_benchmark.py) as the standard entry point for benchmarking; it handles binary detection, argument formatting, and log generation across Windows and Unix platforms.
- The underlying performance derives from [`src/ggml-bitnet-mad.cpp`](https://github.com/microsoft/BitNet/blob/main/src/ggml-bitnet-mad.cpp) (MAD kernels) and [`src/ggml-bitnet-lut.cpp`](https://github.com/microsoft/BitNet/blob/main/src/ggml-bitnet-lut.cpp) (LUT kernels), which implement 1-bit quantized inference.
- Compile the project with `-DCMAKE_BUILD_TYPE=Release` to obtain production-representative throughput numbers.
- Tokens-per-second (tps) is the key metric; compare results across quantization types (`i2_s`, `tl1`, `tl2`) and thread counts to characterize hardware efficiency.
- For single-request latency, use [`run_inference.py`](https://github.com/microsoft/BitNet/blob/main/run_inference.py) instead of the bulk throughput benchmark.

## Frequently Asked Questions

### What is the difference between [`e2e_benchmark.py`](https://github.com/microsoft/BitNet/blob/main/e2e_benchmark.py) and `llama-bench`?

[`e2e_benchmark.py`](https://github.com/microsoft/BitNet/blob/main/e2e_benchmark.py) is a Python convenience wrapper located in [`utils/e2e_benchmark.py`](https://github.com/microsoft/BitNet/blob/main/utils/e2e_benchmark.py) that automates the discovery and invocation of the `llama-bench` binary (located in `build/bin/`). The wrapper standardizes cross-platform execution, manages log file creation, and ensures consistent argument ordering, while `llama-bench` is the raw C++ executable that performs the actual token generation timing.

### Which quantization type provides the fastest inference?

The **tl1** and **tl2** ternary (1.58-bit) quantization schemes generally deliver the highest tps on modern CPUs because they minimize memory bandwidth pressure compared to `i2_s` (2-bit) or higher-precision formats. However, the optimal choice depends on your specific CPU architecture; consult the performance charts in `src/assets/performance.png` to compare speedups on x86 versus ARM platforms.

### How do I benchmark single-prompt latency instead of bulk throughput?

For latency-sensitive applications, use [`run_inference.py`](https://github.com/microsoft/BitNet/blob/main/run_inference.py) with the `time` command to measure wall-clock duration for a single generation request. This approach mimics real-world API usage better than [`e2e_benchmark.py`](https://github.com/microsoft/BitNet/blob/main/e2e_benchmark.py), which is optimized for sustained throughput measurement across multiple iterations.

### Why are my benchmark results slower than the official performance tables?

Ensure you have compiled the project in **Release mode** (`-DCMAKE_BUILD_TYPE=Release`). Debug builds disable vectorization in [`src/ggml-bitnet-mad.cpp`](https://github.com/microsoft/BitNet/blob/main/src/ggml-bitnet-mad.cpp) and [`src/ggml-bitnet-lut.cpp`](https://github.com/microsoft/BitNet/blob/main/src/ggml-bitnet-lut.cpp), resulting in significantly lower tps. Additionally, verify that thermal throttling is not occurring and that you are using the same thread count (`-t`) and model size referenced in the official benchmarks.