# How to Optimize BitNet Inference Performance by Tuning gemm-config.h Parameters

> Boost BitNet inference speed by tuning gemm-config.h parameters like ROW_BLOCK_SIZE and COL_BLOCK_SIZE. Learn manual or automatic optimization for your CPU.

- Repository: [Microsoft/BitNet](https://github.com/microsoft/BitNet)
- Tags: performance
- Published: 2026-03-13

---

**Optimize BitNet inference performance by adjusting the `ROW_BLOCK_SIZE`, `COL_BLOCK_SIZE`, and `PARALLEL_SIZE` macros in [`include/gemm-config.h`](https://github.com/microsoft/BitNet/blob/main/include/gemm-config.h), either manually for specific hardware or automatically using [`utils/tune_gemm_config.py`](https://github.com/microsoft/BitNet/blob/main/utils/tune_gemm_config.py) to benchmark configurations and select the optimal block sizes for your CPU architecture.**

BitNet achieves efficient 1-bit inference through highly optimized GEMM (general matrix multiply) kernels that rely on compile-time constants defined in [[`include/gemm-config.h`](https://github.com/microsoft/BitNet/blob/main/include/gemm-config.h)](https://github.com/microsoft/BitNet/blob/main/include/gemm-config.h). Fine-tuning these parameters to match your specific CPU cache hierarchy and SIMD capabilities can yield **10–30% improvements** in tokens-per-second throughput compared to the default architecture-specific presets.

## Understanding the gemm-config.h Parameters

The file [[`include/gemm-config.h`](https://github.com/microsoft/BitNet/blob/main/include/gemm-config.h)](https://github.com/microsoft/BitNet/blob/main/include/gemm-config.h) defines three critical macros that control the inner loops of the GEMM kernels implemented in [[`src/ggml-bitnet-mad.cpp`](https://github.com/microsoft/BitNet/blob/main/src/ggml-bitnet-mad.cpp)](https://github.com/microsoft/BitNet/blob/main/src/ggml-bitnet-mad.cpp):

### ROW_BLOCK_SIZE

This macro defines how many rows of the **A** matrix are processed in a single inner-loop block. Larger values increase cache reuse for the **A** matrix but risk overflowing L1/L2 caches, which causes memory stalls. The default value varies by architecture: **4** for x86 (AVX/AVX2/AVX-512) and **8** for ARM NEON with DOTPROD.

### COL_BLOCK_SIZE

This parameter sets the number of columns processed per block for the **B** matrix. Increasing this value improves vector-width utilization (critical for AVX-512), but raises the temporary buffer size requirements. Defaults are **128** for x86 and **256** for ARM NEON with DOTPROD to exploit wider SIMD registers.

### PARALLEL_SIZE

This controls the width of the inner-parallel loop that is unrolled and executed simultaneously. Higher values provide more SIMD work per iteration but increase register pressure. The optimal value depends on the number of available CPU registers and should be tuned alongside `ROW_BLOCK_SIZE` and `COL_BLOCK_SIZE`.

### ACT_PARALLEL (Optional)

When defined, this macro compiles the parallel-loop implementation (`ggml_vec_dot_i2_i8_s_Nx1`) in [[`src/ggml-bitnet-mad.cpp`](https://github.com/microsoft/BitNet/blob/main/src/ggml-bitnet-mad.cpp)](https://github.com/microsoft/BitNet/blob/main/src/ggml-bitnet-mad.cpp). When undefined, the compiler uses the sequential version (`ggml_vec_dot_i2_i8_s_1xN`). The fallback path (no AVX, no DOTPROD) uses larger row blocks (128) and smaller column blocks (32) when `ACT_PARALLEL` is disabled.

## Architecture-Specific Defaults

The header contains conditional blocks that select sensible defaults based on the detected instruction set architecture:

- **x86 (AVX/AVX2/AVX-512)**: Uses `ROW_BLOCK_SIZE 4` and `COL_BLOCK_SIZE 128` because wide SIMD registers handle many columns per load efficiently.
- **ARM NEON with DOTPROD**: Uses `ROW_BLOCK_SIZE 8` and `COL_BLOCK_SIZE 256` to maximize utilization of the dot-product instruction.
- **Fallback**: Uses NEON-DOTPROD values when `ACT_PARALLEL` is defined, otherwise falls back to larger row blocks (128) and smaller column blocks (32).

These defaults serve as starting points, but real-world performance depends on your specific CPU cache hierarchy, core count, and model size.

## Automated Tuning with tune_gemm_config.py

Microsoft provides [[`utils/tune_gemm_config.py`](https://github.com/microsoft/BitNet/blob/main/utils/tune_gemm_config.py)](https://github.com/microsoft/BitNet/blob/main/utils/tune_gemm_config.py) to automate the search for optimal parameters. The script performs the following workflow:

1. **Backup**: Creates a timestamped copy of [`gemm-config.h`](https://github.com/microsoft/BitNet/blob/main/gemm-config.h)
2. **Grid Search**: Generates candidate configurations across `ROW_BLOCK_SIZE`, `COL_BLOCK_SIZE`, and `PARALLEL_SIZE`
3. **Compilation**: Rebuilds BitNet with `cmake --build … --target llama-bench` for each candidate
4. **Benchmarking**: Runs `llama-bench` with a representative prompt (`-p 128`) and extracts **pp128 throughput** (tokens/sec)
5. **Logging**: Writes results to `stats/tuning_results_<timestamp>.csv`
6. **Selection**: Identifies the highest throughput configuration
7. **Application**: Overwrites [`gemm-config.h`](https://github.com/microsoft/BitNet/blob/main/gemm-config.h) and rebuilds the project with optimal values

Run the tuner from the repository root:

```bash
python utils/tune_gemm_config.py \
    --config include/gemm-config.h \
    --model models/BitNet-b1.58-2B-4T/ggml-model-i2_s-embed-q6_k.gguf \
    --threads 16 \
    --quick

```

The `--quick` flag uses a reduced configuration grid for faster iteration. Typical output appears as:

```

📦 Backing up current config to gemm-config.h.backup_20240313_101200
🚀 Starting tuning process with 5 configurations
...
✅ PP128: 512.34 ± 10.12 t/s   <-- best so far
...
🏆 BEST CONFIGURATION FOUND!
Configuration: ACT_ON_R4_C128_P4
ACT_PARALLEL: True
ROW_BLOCK_SIZE: 4
COL_BLOCK_SIZE: 128
PARALLEL_SIZE: 4
PP128 Throughput: 523.67 ± 9.45 t/s

```

Answering `y` when prompted applies the configuration to [`include/gemm-config.h`](https://github.com/microsoft/BitNet/blob/main/include/gemm-config.h) and triggers a final rebuild.

## Manual Tuning Workflow

For hardware-specific optimization or when experimenting outside the automated grid, edit [[`include/gemm-config.h`](https://github.com/microsoft/BitNet/blob/main/include/gemm-config.h)](https://github.com/microsoft/BitNet/blob/main/include/gemm-config.h) directly:

```c
#define ACT_PARALLEL
#define ROW_BLOCK_SIZE 8          // Increase row processing
#define COL_BLOCK_SIZE 256        // Exploit AVX-512 width
#define PARALLEL_SIZE 8           // Higher parallelism per loop

```

After saving changes, perform a clean rebuild:

```bash
cmake -B build -S . -DCMAKE_BUILD_TYPE=Release
cmake --build build --target llama-bench

```

Validate the configuration by benchmarking:

```bash
build/bin/llama-bench \
    -m models/BitNet-b1.58-2B-4T/ggml-model-i2_s-embed-q6_k.gguf \
    -p 128 -n 0 -t 16 -ngl 0

```

Compare the `pp128` column (reported as `| pp128 | 523.67 ± 9.45 |`) against your baseline to confirm speedup.

## Custom Configuration Lists

To test specific combinations without manual editing, create a Python file defining your candidate configurations:

```python

# custom_configs.py

configs = [
    {"act_parallel": True, "row_block_size": 8, "col_block_size": 256, "parallel_size": 8},
    {"act_parallel": False, "row_block_size": 128, "col_block_size": 32, "parallel_size": 4},
]

```

Then pass it to the tuner:

```bash
python utils/tune_gemm_config.py \
    --config include/gemm-config.h \
    --model models/BitNet-b1.58-2B-4T/ggml-model-i2_s-embed-q6_k.gguf \
    --threads 8 \
    --custom < custom_configs.py

```

## Summary

- **Three core macros** control BitNet GEMM performance: `ROW_BLOCK_SIZE`, `COL_BLOCK_SIZE`, and `PARALLEL_SIZE`, defined in [[`include/gemm-config.h`](https://github.com/microsoft/BitNet/blob/main/include/gemm-config.h)](https://github.com/microsoft/BitNet/blob/main/include/gemm-config.h)
- **Architecture defaults** differ between x86 (4×128) and ARM NEON (8×256) to match SIMD register widths
- **Automated tuning** via [`utils/tune_gemm_config.py`](https://github.com/microsoft/BitNet/blob/main/utils/tune_gemm_config.py) can improve throughput by 10–30% by grid-searching configurations and benchmarking with `llama-bench`
- **Manual tuning** requires editing the header, rebuilding with CMake, and validating through the `pp128` metric in `llama-bench`
- **ACT_PARALLEL** toggles between parallel (`ggml_vec_dot_i2_i8_s_Nx1`) and sequential (`ggml_vec_dot_i2_i8_s_1xN`) kernel implementations

## Frequently Asked Questions

### What is the difference between ROW_BLOCK_SIZE and COL_BLOCK_SIZE?

`ROW_BLOCK_SIZE` controls how many rows of the **A** matrix are processed per block, affecting cache locality for the left-hand matrix. `COL_BLOCK_SIZE` controls columns of the **B** matrix processed per block, impacting vector-width utilization (SIMD efficiency). According to the BitNet source code, x86 architectures typically favor smaller row blocks (4) and wider column blocks (128), while ARM NEON benefits from larger values in both dimensions (8 and 256).

### How do I know if ACT_PARALLEL should be enabled?

Enable `ACT_PARALLEL` (defined in [`include/gemm-config.h`](https://github.com/microsoft/BitNet/blob/main/include/gemm-config.h)) when your CPU has sufficient registers to handle the parallel-loop implementation `ggml_vec_dot_i2_i8_s_Nx1`. This version unrolls the inner loop wider than the sequential `ggml_vec_dot_i2_i8_s_1xN` variant. The [`tune_gemm_config.py`](https://github.com/microsoft/BitNet/blob/main/tune_gemm_config.py) script automatically tests both variants and reports which yields higher pp128 throughput for your specific hardware.

### Can I tune these parameters for non-AVX CPUs?

Yes. The fallback path in [`gemm-config.h`](https://github.com/microsoft/BitNet/blob/main/gemm-config.h) supports generic CPUs without AVX or NEON DOTPROD. When `ACT_PARALLEL` is undefined, the code path uses `ROW_BLOCK_SIZE 128` and `COL_BLOCK_SIZE 32`. You can manually adjust these values or use the tuner script, which will compile and benchmark configurations regardless of the detected ISA, though the performance gains may differ from SIMD-enabled architectures.

### Why does the tuning script use llama-bench instead of run_inference.py?

`llama-bench` provides deterministic, low-overhead measurement of pure inference throughput (specifically the pp128 metric) without the additional latency of Python bindings or server overhead present in [`run_inference.py`](https://github.com/microsoft/BitNet/blob/main/run_inference.py) or [`run_inference_server.py`](https://github.com/microsoft/BitNet/blob/main/run_inference_server.py). The tuner requires consistent, reproducible metrics to compare compile-time configurations, and the C++ benchmark binary offers lower variance than the Python wrappers.