How to Use Pretrained Kernel Parameters for BitNet Optimization

Use the --use-pretuned flag with setup_env.py to automatically select model-specific GPU kernel tiling configurations (BM, BK, bmm) that eliminate runtime transposes and maximize throughput.

Microsoft's BitNet repository provides pretrained kernel parameters that optimize inference by pre-tuning weight matrix tiling sizes for specific GPU architectures. When you use these parameters, the system automatically configures exact tiling dimensions (BM, BK, and micro-tile bmm) during model conversion, ensuring the quantized weights align perfectly with the compiled kernels. This approach removes the need for manual profiling and delivers peak performance out-of-the-box for supported models like bitnet_b1_58-large and Llama-3 variants.

How Pretrained Kernel Parameters Work in BitNet

BitNet's inference performance depends on how weight matrices are tiled and quantized for GPU execution. The repository ships manually profiled configurations stored in preset_kernels/<model_name>/ directories, containing:

When enabled, setup_env.py copies these files into the include/ directory. During model conversion, utils/convert-hf-to-gguf-bitnet.py (lines 93-100) reads the configuration via ConfigParser and reshapes weights using preprocess_weights_tl1 to match the expected tiling, eliminating runtime transpose operations.

Enabling Pretuned Kernels During Setup

To activate pretrained kernel parameters, pass the --use-pretuned flag when running the environment setup script. The script automatically detects your model architecture and selects the appropriate TL-1 or TL-2 configuration.

python setup_env.py \
    --hf-repo microsoft/bitnet-b1-58-large \
    --quant-type tl1 \
    --use-pretuned

This command performs three critical operations defined in setup_env.py:

  1. Detects the model architecture and locates the matching folder in preset_kernels/
  2. Copies bitnet-lut-kernels-tl1.h to include/bitnet-lut-kernels.h
  3. Copies kernel_config_tl1.ini to include/kernel_config.ini

Converting Models with Kernel-Aware Weight Packing

After setup, convert your HuggingFace checkpoint using the pretuned configuration. The conversion script reads include/kernel_config.ini to determine exact reshaping parameters for each weight matrix.

python utils/convert-hf-to-gguf-bitnet.py \
    --hf-repo microsoft/bitnet-b1-58-large \
    --output-dir models/bitnet_b1_58-large

According to the source code in utils/convert-hf-to-gguf-bitnet.py (lines 93-99), the script iterates through kernel configuration sections to find entries where m and k match the weight dimensions, extracting BM, BK, and bmm values. These drive the preprocess_weights_tl1 pipeline that packs tensors into the exact format expected by the kernels.

Building and Running Optimized Inference

With kernel sources and configurations in place, setup_env.py invokes CMake to compile the inference binary using the pretuned tile sizes. The runtime (run_inference.py) then loads these pre-packed weights and calls the compiled kernels without additional overhead.

python run_inference.py \
    --model-dir models/bitnet_b1_58-large \
    --quant-type tl1 \
    --prompt "What is the capital of France?"

Understanding Kernel Configuration File Structure

The .ini files in preset_kernels/<model>/ define tiling parameters for specific matrix shapes. Each section represents a kernel block:

[Kernels_0]
m = 1536
k = 4096
bm = 256
bk = 96
bmm = 32
  • m and k: Input matrix dimensions (rows and columns)
  • bm (BM): Block size in the M dimension
  • bk (BK): Block size in the K dimension
  • bmm: Micro-tile size for the matrix multiplication

These values are generated by utils/codegen_tl1.py and utils/codegen_tl2.py, which create the kernel headers based on the model's shape dictionary.

Loading Parameters Manually in Python

You can programmatically access pretuned parameters using the same ConfigParser logic found in the conversion utilities:

from configparser import ConfigParser
import os

def get_tile_params(m, k, config_path="include/kernel_config.ini"):
    cfg = ConfigParser()
    cfg.read(config_path)
    
    for section in cfg.sections():
        if int(cfg[section]["m"]) == m and int(cfg[section]["k"]) == k:
            return {
                "BM": int(cfg[section]["bm"]),
                "BK": int(cfg[section]["bk"]),
                "bmm": int(cfg[section]["bmm"]),
            }
    raise ValueError(f"No pretuned tile for shape ({m}, {k})")

# Retrieve parameters for a 1536×4096 weight matrix

params = get_tile_params(1536, 4096)
print(params)  # Output: {'BM': 256, 'BK': 96, 'bmm': 32}

This matches the internal implementation used by preprocess_weights_tl1 to ensure weight tensors align with kernel expectations.

Summary

  • Pretrained kernel parameters eliminate manual GPU kernel tuning by providing model-specific tile sizes (BM, BK, bmm) optimized for specific architectures.
  • Activation method: Use the --use-pretuned flag with setup_env.py to automatically copy configuration files from preset_kernels/<model>/ to include/.
  • Weight alignment: The convert-hf-to-gguf-bitnet.py script reads include/kernel_config.ini via ConfigParser (lines 93-100) and reshapes weights using preprocess_weights_tl1 to match kernel tiling.
  • File locations: Configurations reside in preset_kernels/ (e.g., bitnet_b1_58-large/kernel_config_tl1.ini) and are generated by utils/codegen_tl1.py and codegen_tl2.py.
  • Performance impact: Proper tiling ensures the generated kernels consume tensors without runtime transposes, delivering maximum throughput during inference.

Frequently Asked Questions

What models support pretrained kernel parameters?

The preset_kernels/ directory contains optimized configurations for bitnet_b1_58-large and various Llama-3 variants. Each model folder includes both TL-1 and TL-2 quantization presets. Check the repository's preset_kernels/ directory for the current list of supported architectures.

How do BM, BK, and bmm parameters affect performance?

BM (block size M), BK (block size K), and bmm (micro-tile) determine how weight matrices are partitioned for GPU computation. These values are manually profiled to match specific GPU cache sizes and memory bandwidth characteristics. Incorrect values cause memory bank conflicts and reduced parallelism, while pretuned values maximize occupancy and minimize data movement.

Can I use pretuned kernels with custom model architectures?

No, pretuned kernels are model-specific because the tiling parameters depend on exact weight matrix dimensions. For custom architectures, you must either generate new kernel parameters using utils/codegen_tl1.py or codegen_tl2.py, or use the default auto-tuning path without --use-pretuned, though this may yield suboptimal performance compared to the manually profiled presets.

Where are the kernel source files located after setup?

When using --use-pretuned, setup_env.py copies bitnet-lut-kernels-tl1.h (or TL-2 variant) to include/bitnet-lut-kernels.h and kernel_config_tl1.ini to include/kernel_config.ini. The original files remain in preset_kernels/<model_name>/, preserving the source templates for reference.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →