# How to Use Pretrained Kernel Parameters for BitNet Optimization

> Optimize BitNet models using pretrained kernel parameters. Leverage the use pretuned flag in setup_env.py to boost GPU throughput and eliminate runtime transposes for faster performance.

- Repository: [Microsoft/BitNet](https://github.com/microsoft/BitNet)
- Tags: how-to-guide
- Published: 2026-03-13

---

**Use the `--use-pretuned` flag with [`setup_env.py`](https://github.com/microsoft/BitNet/blob/main/setup_env.py) to automatically select model-specific GPU kernel tiling configurations (BM, BK, bmm) that eliminate runtime transposes and maximize throughput.**

Microsoft's BitNet repository provides pretrained kernel parameters that optimize inference by pre-tuning weight matrix tiling sizes for specific GPU architectures. When you use these parameters, the system automatically configures exact tiling dimensions (BM, BK, and micro-tile bmm) during model conversion, ensuring the quantized weights align perfectly with the compiled kernels. This approach removes the need for manual profiling and delivers peak performance out-of-the-box for supported models like `bitnet_b1_58-large` and Llama-3 variants.

## How Pretrained Kernel Parameters Work in BitNet

BitNet's inference performance depends on how weight matrices are tiled and quantized for GPU execution. The repository ships manually profiled configurations stored in `preset_kernels/<model_name>/` directories, containing:

- [`kernel_config_tl1.ini`](https://github.com/microsoft/BitNet/blob/main/kernel_config_tl1.ini) and [`kernel_config_tl2.ini`](https://github.com/microsoft/BitNet/blob/main/kernel_config_tl2.ini): Configuration files mapping weight shapes (m, k) to optimal tile sizes
- [`bitnet-lut-kernels-tl1.h`](https://github.com/microsoft/BitNet/blob/main/bitnet-lut-kernels-tl1.h) and [`bitnet-lut-kernels-tl2.h`](https://github.com/microsoft/BitNet/blob/main/bitnet-lut-kernels-tl2.h): Pre-generated kernel headers compiled with these specific parameters

When enabled, [`setup_env.py`](https://github.com/microsoft/BitNet/blob/main/setup_env.py) copies these files into the `include/` directory. During model conversion, [`utils/convert-hf-to-gguf-bitnet.py`](https://github.com/microsoft/BitNet/blob/main/utils/convert-hf-to-gguf-bitnet.py) (lines 93-100) reads the configuration via `ConfigParser` and reshapes weights using `preprocess_weights_tl1` to match the expected tiling, eliminating runtime transpose operations.

## Enabling Pretuned Kernels During Setup

To activate pretrained kernel parameters, pass the `--use-pretuned` flag when running the environment setup script. The script automatically detects your model architecture and selects the appropriate TL-1 or TL-2 configuration.

```bash
python setup_env.py \
    --hf-repo microsoft/bitnet-b1-58-large \
    --quant-type tl1 \
    --use-pretuned

```

This command performs three critical operations defined in [`setup_env.py`](https://github.com/microsoft/BitNet/blob/main/setup_env.py):

1. Detects the model architecture and locates the matching folder in `preset_kernels/`
2. Copies [`bitnet-lut-kernels-tl1.h`](https://github.com/microsoft/BitNet/blob/main/bitnet-lut-kernels-tl1.h) to [`include/bitnet-lut-kernels.h`](https://github.com/microsoft/BitNet/blob/main/include/bitnet-lut-kernels.h)
3. Copies [`kernel_config_tl1.ini`](https://github.com/microsoft/BitNet/blob/main/kernel_config_tl1.ini) to [`include/kernel_config.ini`](https://github.com/microsoft/BitNet/blob/main/include/kernel_config.ini)

## Converting Models with Kernel-Aware Weight Packing

After setup, convert your HuggingFace checkpoint using the pretuned configuration. The conversion script reads [`include/kernel_config.ini`](https://github.com/microsoft/BitNet/blob/main/include/kernel_config.ini) to determine exact reshaping parameters for each weight matrix.

```bash
python utils/convert-hf-to-gguf-bitnet.py \
    --hf-repo microsoft/bitnet-b1-58-large \
    --output-dir models/bitnet_b1_58-large

```

According to the source code in [`utils/convert-hf-to-gguf-bitnet.py`](https://github.com/microsoft/BitNet/blob/main/utils/convert-hf-to-gguf-bitnet.py) (lines 93-99), the script iterates through kernel configuration sections to find entries where `m` and `k` match the weight dimensions, extracting `BM`, `BK`, and `bmm` values. These drive the `preprocess_weights_tl1` pipeline that packs tensors into the exact format expected by the kernels.

## Building and Running Optimized Inference

With kernel sources and configurations in place, [`setup_env.py`](https://github.com/microsoft/BitNet/blob/main/setup_env.py) invokes CMake to compile the inference binary using the pretuned tile sizes. The runtime ([`run_inference.py`](https://github.com/microsoft/BitNet/blob/main/run_inference.py)) then loads these pre-packed weights and calls the compiled kernels without additional overhead.

```bash
python run_inference.py \
    --model-dir models/bitnet_b1_58-large \
    --quant-type tl1 \
    --prompt "What is the capital of France?"

```

## Understanding Kernel Configuration File Structure

The `.ini` files in `preset_kernels/<model>/` define tiling parameters for specific matrix shapes. Each section represents a kernel block:

```ini
[Kernels_0]
m = 1536
k = 4096
bm = 256
bk = 96
bmm = 32

```

- **m** and **k**: Input matrix dimensions (rows and columns)
- **bm** (BM): Block size in the M dimension
- **bk** (BK): Block size in the K dimension
- **bmm**: Micro-tile size for the matrix multiplication

These values are generated by [`utils/codegen_tl1.py`](https://github.com/microsoft/BitNet/blob/main/utils/codegen_tl1.py) and [`utils/codegen_tl2.py`](https://github.com/microsoft/BitNet/blob/main/utils/codegen_tl2.py), which create the kernel headers based on the model's shape dictionary.

## Loading Parameters Manually in Python

You can programmatically access pretuned parameters using the same `ConfigParser` logic found in the conversion utilities:

```python
from configparser import ConfigParser
import os

def get_tile_params(m, k, config_path="include/kernel_config.ini"):
    cfg = ConfigParser()
    cfg.read(config_path)
    
    for section in cfg.sections():
        if int(cfg[section]["m"]) == m and int(cfg[section]["k"]) == k:
            return {
                "BM": int(cfg[section]["bm"]),
                "BK": int(cfg[section]["bk"]),
                "bmm": int(cfg[section]["bmm"]),
            }
    raise ValueError(f"No pretuned tile for shape ({m}, {k})")

# Retrieve parameters for a 1536×4096 weight matrix

params = get_tile_params(1536, 4096)
print(params)  # Output: {'BM': 256, 'BK': 96, 'bmm': 32}

```

This matches the internal implementation used by `preprocess_weights_tl1` to ensure weight tensors align with kernel expectations.

## Summary

- **Pretrained kernel parameters** eliminate manual GPU kernel tuning by providing model-specific tile sizes (BM, BK, bmm) optimized for specific architectures.
- **Activation method**: Use the `--use-pretuned` flag with [`setup_env.py`](https://github.com/microsoft/BitNet/blob/main/setup_env.py) to automatically copy configuration files from `preset_kernels/<model>/` to `include/`.
- **Weight alignment**: The [`convert-hf-to-gguf-bitnet.py`](https://github.com/microsoft/BitNet/blob/main/convert-hf-to-gguf-bitnet.py) script reads [`include/kernel_config.ini`](https://github.com/microsoft/BitNet/blob/main/include/kernel_config.ini) via `ConfigParser` (lines 93-100) and reshapes weights using `preprocess_weights_tl1` to match kernel tiling.
- **File locations**: Configurations reside in `preset_kernels/` (e.g., [`bitnet_b1_58-large/kernel_config_tl1.ini`](https://github.com/microsoft/BitNet/blob/main/bitnet_b1_58-large/kernel_config_tl1.ini)) and are generated by [`utils/codegen_tl1.py`](https://github.com/microsoft/BitNet/blob/main/utils/codegen_tl1.py) and [`codegen_tl2.py`](https://github.com/microsoft/BitNet/blob/main/codegen_tl2.py).
- **Performance impact**: Proper tiling ensures the generated kernels consume tensors without runtime transposes, delivering maximum throughput during inference.

## Frequently Asked Questions

### What models support pretrained kernel parameters?

The `preset_kernels/` directory contains optimized configurations for `bitnet_b1_58-large` and various Llama-3 variants. Each model folder includes both TL-1 and TL-2 quantization presets. Check the repository's `preset_kernels/` directory for the current list of supported architectures.

### How do BM, BK, and bmm parameters affect performance?

**BM** (block size M), **BK** (block size K), and **bmm** (micro-tile) determine how weight matrices are partitioned for GPU computation. These values are manually profiled to match specific GPU cache sizes and memory bandwidth characteristics. Incorrect values cause memory bank conflicts and reduced parallelism, while pretuned values maximize occupancy and minimize data movement.

### Can I use pretuned kernels with custom model architectures?

No, pretuned kernels are model-specific because the tiling parameters depend on exact weight matrix dimensions. For custom architectures, you must either generate new kernel parameters using [`utils/codegen_tl1.py`](https://github.com/microsoft/BitNet/blob/main/utils/codegen_tl1.py) or [`codegen_tl2.py`](https://github.com/microsoft/BitNet/blob/main/codegen_tl2.py), or use the default auto-tuning path without `--use-pretuned`, though this may yield suboptimal performance compared to the manually profiled presets.

### Where are the kernel source files located after setup?

When using `--use-pretuned`, [`setup_env.py`](https://github.com/microsoft/BitNet/blob/main/setup_env.py) copies [`bitnet-lut-kernels-tl1.h`](https://github.com/microsoft/BitNet/blob/main/bitnet-lut-kernels-tl1.h) (or TL-2 variant) to [`include/bitnet-lut-kernels.h`](https://github.com/microsoft/BitNet/blob/main/include/bitnet-lut-kernels.h) and [`kernel_config_tl1.ini`](https://github.com/microsoft/BitNet/blob/main/kernel_config_tl1.ini) to [`include/kernel_config.ini`](https://github.com/microsoft/BitNet/blob/main/include/kernel_config.ini). The original files remain in `preset_kernels/<model_name>/`, preserving the source templates for reference.