How to Optimize BitNet Inference Performance by Tuning gemm-config.h Parameters

Optimize BitNet inference performance by adjusting the ROW_BLOCK_SIZE, COL_BLOCK_SIZE, and PARALLEL_SIZE macros in include/gemm-config.h, either manually for specific hardware or automatically using utils/tune_gemm_config.py to benchmark configurations and select the optimal block sizes for your CPU architecture.

BitNet achieves efficient 1-bit inference through highly optimized GEMM (general matrix multiply) kernels that rely on compile-time constants defined in [include/gemm-config.h](https://github.com/microsoft/BitNet/blob/main/include/gemm-config.h). Fine-tuning these parameters to match your specific CPU cache hierarchy and SIMD capabilities can yield 10–30% improvements in tokens-per-second throughput compared to the default architecture-specific presets.

Understanding the gemm-config.h Parameters

The file [include/gemm-config.h](https://github.com/microsoft/BitNet/blob/main/include/gemm-config.h) defines three critical macros that control the inner loops of the GEMM kernels implemented in [src/ggml-bitnet-mad.cpp](https://github.com/microsoft/BitNet/blob/main/src/ggml-bitnet-mad.cpp):

ROW_BLOCK_SIZE

This macro defines how many rows of the A matrix are processed in a single inner-loop block. Larger values increase cache reuse for the A matrix but risk overflowing L1/L2 caches, which causes memory stalls. The default value varies by architecture: 4 for x86 (AVX/AVX2/AVX-512) and 8 for ARM NEON with DOTPROD.

COL_BLOCK_SIZE

This parameter sets the number of columns processed per block for the B matrix. Increasing this value improves vector-width utilization (critical for AVX-512), but raises the temporary buffer size requirements. Defaults are 128 for x86 and 256 for ARM NEON with DOTPROD to exploit wider SIMD registers.

PARALLEL_SIZE

This controls the width of the inner-parallel loop that is unrolled and executed simultaneously. Higher values provide more SIMD work per iteration but increase register pressure. The optimal value depends on the number of available CPU registers and should be tuned alongside ROW_BLOCK_SIZE and COL_BLOCK_SIZE.

ACT_PARALLEL (Optional)

When defined, this macro compiles the parallel-loop implementation (ggml_vec_dot_i2_i8_s_Nx1) in [src/ggml-bitnet-mad.cpp](https://github.com/microsoft/BitNet/blob/main/src/ggml-bitnet-mad.cpp). When undefined, the compiler uses the sequential version (ggml_vec_dot_i2_i8_s_1xN). The fallback path (no AVX, no DOTPROD) uses larger row blocks (128) and smaller column blocks (32) when ACT_PARALLEL is disabled.

Architecture-Specific Defaults

The header contains conditional blocks that select sensible defaults based on the detected instruction set architecture:

  • x86 (AVX/AVX2/AVX-512): Uses ROW_BLOCK_SIZE 4 and COL_BLOCK_SIZE 128 because wide SIMD registers handle many columns per load efficiently.
  • ARM NEON with DOTPROD: Uses ROW_BLOCK_SIZE 8 and COL_BLOCK_SIZE 256 to maximize utilization of the dot-product instruction.
  • Fallback: Uses NEON-DOTPROD values when ACT_PARALLEL is defined, otherwise falls back to larger row blocks (128) and smaller column blocks (32).

These defaults serve as starting points, but real-world performance depends on your specific CPU cache hierarchy, core count, and model size.

Automated Tuning with tune_gemm_config.py

Microsoft provides [utils/tune_gemm_config.py](https://github.com/microsoft/BitNet/blob/main/utils/tune_gemm_config.py) to automate the search for optimal parameters. The script performs the following workflow:

  1. Backup: Creates a timestamped copy of gemm-config.h
  2. Grid Search: Generates candidate configurations across ROW_BLOCK_SIZE, COL_BLOCK_SIZE, and PARALLEL_SIZE
  3. Compilation: Rebuilds BitNet with cmake --build … --target llama-bench for each candidate
  4. Benchmarking: Runs llama-bench with a representative prompt (-p 128) and extracts pp128 throughput (tokens/sec)
  5. Logging: Writes results to stats/tuning_results_<timestamp>.csv
  6. Selection: Identifies the highest throughput configuration
  7. Application: Overwrites gemm-config.h and rebuilds the project with optimal values

Run the tuner from the repository root:

python utils/tune_gemm_config.py \
    --config include/gemm-config.h \
    --model models/BitNet-b1.58-2B-4T/ggml-model-i2_s-embed-q6_k.gguf \
    --threads 16 \
    --quick

The --quick flag uses a reduced configuration grid for faster iteration. Typical output appears as:


📦 Backing up current config to gemm-config.h.backup_20240313_101200
🚀 Starting tuning process with 5 configurations
...
✅ PP128: 512.34 ± 10.12 t/s   <-- best so far
...
🏆 BEST CONFIGURATION FOUND!
Configuration: ACT_ON_R4_C128_P4
ACT_PARALLEL: True
ROW_BLOCK_SIZE: 4
COL_BLOCK_SIZE: 128
PARALLEL_SIZE: 4
PP128 Throughput: 523.67 ± 9.45 t/s

Answering y when prompted applies the configuration to include/gemm-config.h and triggers a final rebuild.

Manual Tuning Workflow

For hardware-specific optimization or when experimenting outside the automated grid, edit [include/gemm-config.h](https://github.com/microsoft/BitNet/blob/main/include/gemm-config.h) directly:

#define ACT_PARALLEL
#define ROW_BLOCK_SIZE 8          // Increase row processing
#define COL_BLOCK_SIZE 256        // Exploit AVX-512 width
#define PARALLEL_SIZE 8           // Higher parallelism per loop

After saving changes, perform a clean rebuild:

cmake -B build -S . -DCMAKE_BUILD_TYPE=Release
cmake --build build --target llama-bench

Validate the configuration by benchmarking:

build/bin/llama-bench \
    -m models/BitNet-b1.58-2B-4T/ggml-model-i2_s-embed-q6_k.gguf \
    -p 128 -n 0 -t 16 -ngl 0

Compare the pp128 column (reported as | pp128 | 523.67 ± 9.45 |) against your baseline to confirm speedup.

Custom Configuration Lists

To test specific combinations without manual editing, create a Python file defining your candidate configurations:


# custom_configs.py

configs = [
    {"act_parallel": True, "row_block_size": 8, "col_block_size": 256, "parallel_size": 8},
    {"act_parallel": False, "row_block_size": 128, "col_block_size": 32, "parallel_size": 4},
]

Then pass it to the tuner:

python utils/tune_gemm_config.py \
    --config include/gemm-config.h \
    --model models/BitNet-b1.58-2B-4T/ggml-model-i2_s-embed-q6_k.gguf \
    --threads 8 \
    --custom < custom_configs.py

Summary

  • Three core macros control BitNet GEMM performance: ROW_BLOCK_SIZE, COL_BLOCK_SIZE, and PARALLEL_SIZE, defined in [include/gemm-config.h](https://github.com/microsoft/BitNet/blob/main/include/gemm-config.h)
  • Architecture defaults differ between x86 (4×128) and ARM NEON (8×256) to match SIMD register widths
  • Automated tuning via utils/tune_gemm_config.py can improve throughput by 10–30% by grid-searching configurations and benchmarking with llama-bench
  • Manual tuning requires editing the header, rebuilding with CMake, and validating through the pp128 metric in llama-bench
  • ACT_PARALLEL toggles between parallel (ggml_vec_dot_i2_i8_s_Nx1) and sequential (ggml_vec_dot_i2_i8_s_1xN) kernel implementations

Frequently Asked Questions

What is the difference between ROW_BLOCK_SIZE and COL_BLOCK_SIZE?

ROW_BLOCK_SIZE controls how many rows of the A matrix are processed per block, affecting cache locality for the left-hand matrix. COL_BLOCK_SIZE controls columns of the B matrix processed per block, impacting vector-width utilization (SIMD efficiency). According to the BitNet source code, x86 architectures typically favor smaller row blocks (4) and wider column blocks (128), while ARM NEON benefits from larger values in both dimensions (8 and 256).

How do I know if ACT_PARALLEL should be enabled?

Enable ACT_PARALLEL (defined in include/gemm-config.h) when your CPU has sufficient registers to handle the parallel-loop implementation ggml_vec_dot_i2_i8_s_Nx1. This version unrolls the inner loop wider than the sequential ggml_vec_dot_i2_i8_s_1xN variant. The tune_gemm_config.py script automatically tests both variants and reports which yields higher pp128 throughput for your specific hardware.

Can I tune these parameters for non-AVX CPUs?

Yes. The fallback path in gemm-config.h supports generic CPUs without AVX or NEON DOTPROD. When ACT_PARALLEL is undefined, the code path uses ROW_BLOCK_SIZE 128 and COL_BLOCK_SIZE 32. You can manually adjust these values or use the tuner script, which will compile and benchmark configurations regardless of the detected ISA, though the performance gains may differ from SIMD-enabled architectures.

Why does the tuning script use llama-bench instead of run_inference.py?

llama-bench provides deterministic, low-overhead measurement of pure inference throughput (specifically the pp128 metric) without the additional latency of Python bindings or server overhead present in run_inference.py or run_inference_server.py. The tuner requires consistent, reproducible metrics to compare compile-time configurations, and the C++ benchmark binary offers lower variance than the Python wrappers.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →