How to Benchmark BitNet Inference Performance
Use the utils/e2e_benchmark.py script to drive the compiled llama-bench binary, measuring tokens-per-second (tps) across different thread counts, prompt lengths, and model configurations.
Microsoft/BitNet ships with a built-in end-to-end benchmarking system designed to standardize performance measurement across x86 and ARM platforms. To accurately benchmark BitNet inference performance, you will compile the optimized C++ kernels, prepare a quantized GGUF model, and execute the Python wrapper that orchestrates the llama-bench binary with consistent parameters.
Architecture of the Benchmark System
The benchmarking workflow relies on a Python frontend and a high-performance C++ backend. Understanding how these components interact ensures you can troubleshoot issues and customize measurements.
Binary Discovery and Cross-Platform Handling
In utils/e2e_benchmark.py, lines L25-L33 implement automatic discovery of the benchmark executable. The script searches the build/bin directory and automatically selects the correct binary name for your platform—llama-bench.exe on Windows or llama-bench on Unix-like systems—eliminating manual path configuration.
Command Construction and Execution
Lines L36-L45 of utils/e2e_benchmark.py assemble the command-line arguments for the benchmark binary. The script injects the model path (-m), generation length (-n), prompt size (-p), thread count (-t), and a fixed repeat count of five (-r 5) to ensure statistical significance. Lines L8-L18 handle execution via subprocess.run, capturing both stdout and stderr to a timestamped log file in your specified log_dir (defaulting to the current working directory).
Core Inference Kernels
The speed advantages of BitNet stem from specialized 1-bit matrix operations implemented in the C++ source. The Multiply-Add (MAD) kernel lives in src/ggml-bitnet-mad.cpp (lines L1-L20), while the Lookup-Table (LUT) kernel resides in src/ggml-bitnet-lut.cpp (lines L1-L20). These kernels replace traditional floating-point multiplication with bitwise operations, enabling the high throughput measured by the benchmark.
Step-by-Step Benchmarking Workflow
Follow these steps to generate reproducible performance metrics for any BitNet model.
-
Build the project in release mode
Compile the
llama-benchbinary and core libraries with optimizations enabled. Release mode is critical for realistic throughput numbers.git clone --recursive https://github.com/microsoft/BitNet.git cd BitNet mkdir -p build && cd build cmake -DCMAKE_BUILD_TYPE=Release .. && make -j$(nproc) -
Download and prepare a BitNet model
Retrieve a quantized GGUF model from Hugging Face and run the setup script to prepare the environment.
huggingface-cli download microsoft/BitNet-b1.58-2B-4T --local-dir models/BitNet-2B python setup_env.py -md models/BitNet-2B -q i2_s -
Run the end-to-end benchmark
Execute
utils/e2e_benchmark.pywith your model path and desired test parameters. The following example generates 200 tokens from a 256-token prompt using 8 threads:python utils/e2e_benchmark.py \ -m models/BitNet-2B/ggml-model-i2_s.gguf \ -n 200 \ -p 256 \ -t 8The script repeats the test five times and writes raw
llama-benchoutput to a log file namede2e_benchmark_<timestamp>.log. -
Inspect the results
Read the generated log to extract the tokens-per-second (tps) metric. Look for lines containing
tps:followed by the numerical value.cat e2e_benchmark_$(date +%s).logTypical output includes per-iteration timing and an aggregate tps average across the five repeats.
-
(Optional) Measure single-prompt latency
For wall-clock latency measurements of individual requests, use the high-level Python wrapper instead of the throughput benchmark:
time python run_inference.py \ -m models/BitNet-2B/ggml-model-i2_s.gguf \ -p "Explain quantum computing" \ -t 8
Interpreting Benchmark Results
Tokens-per-second (tps) is the primary metric for throughput. Higher values indicate faster inference. When you benchmark BitNet inference performance, compare results across these dimensions:
- Quantization schemes: Test
i2_s(2-bit),tl1, andtl2(ternary 1.58-bit) configurations to evaluate the speed-precision trade-off. Lower bit widths generally yield higher tps on CPU. - Thread scaling: Increase the
-tparameter to measure how well the MAD and LUT kernels utilize multi-core x86 or ARM processors. Reference the performance tables insrc/assets/performance.png(cited inREADME.mdlines L15-L18) for expected scaling curves on specific hardware. - Model size impact: Compare the 2B parameter model against larger variants to quantify the overhead of increased model capacity versus the efficiency of 1-bit compute.
Best Practices for Accurate Results
To avoid skewed measurements when using utils/e2e_benchmark.py, adhere to these guidelines:
- Always build in release mode. Debug builds (
-DCMAKE_BUILD_TYPE=Debug) disable compiler optimizations and yield unrealistically low tps. - Disable CPU frequency scaling. Power-saving features like Intel SpeedStep or AMD Cool'n'Quiet can throttle cores during the benchmark run, creating variance between iterations.
- Account for warm-up effects. The script executes five repeats (
-r 5); discard the first iteration to allow CPU caches and branch predictors to stabilize, then average the remaining four runs. - Isolate the test environment. Close background applications that compete for CPU cycles or memory bandwidth to ensure the
llama-benchbinary has exclusive resource access.
Summary
- Use
utils/e2e_benchmark.pyas the standard entry point for benchmarking; it handles binary detection, argument formatting, and log generation across Windows and Unix platforms. - The underlying performance derives from
src/ggml-bitnet-mad.cpp(MAD kernels) andsrc/ggml-bitnet-lut.cpp(LUT kernels), which implement 1-bit quantized inference. - Compile the project with
-DCMAKE_BUILD_TYPE=Releaseto obtain production-representative throughput numbers. - Tokens-per-second (tps) is the key metric; compare results across quantization types (
i2_s,tl1,tl2) and thread counts to characterize hardware efficiency. - For single-request latency, use
run_inference.pyinstead of the bulk throughput benchmark.
Frequently Asked Questions
What is the difference between e2e_benchmark.py and llama-bench?
e2e_benchmark.py is a Python convenience wrapper located in utils/e2e_benchmark.py that automates the discovery and invocation of the llama-bench binary (located in build/bin/). The wrapper standardizes cross-platform execution, manages log file creation, and ensures consistent argument ordering, while llama-bench is the raw C++ executable that performs the actual token generation timing.
Which quantization type provides the fastest inference?
The tl1 and tl2 ternary (1.58-bit) quantization schemes generally deliver the highest tps on modern CPUs because they minimize memory bandwidth pressure compared to i2_s (2-bit) or higher-precision formats. However, the optimal choice depends on your specific CPU architecture; consult the performance charts in src/assets/performance.png to compare speedups on x86 versus ARM platforms.
How do I benchmark single-prompt latency instead of bulk throughput?
For latency-sensitive applications, use run_inference.py with the time command to measure wall-clock duration for a single generation request. This approach mimics real-world API usage better than e2e_benchmark.py, which is optimized for sustained throughput measurement across multiple iterations.
Why are my benchmark results slower than the official performance tables?
Ensure you have compiled the project in Release mode (-DCMAKE_BUILD_TYPE=Release). Debug builds disable vectorization in src/ggml-bitnet-mad.cpp and src/ggml-bitnet-lut.cpp, resulting in significantly lower tps. Additionally, verify that thermal throttling is not occurring and that you are using the same thread count (-t) and model size referenced in the official benchmarks.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →