How to Use Pretrained Kernel Parameters for BitNet Optimization
Use the --use-pretuned flag with setup_env.py to automatically select model-specific GPU kernel tiling configurations (BM, BK, bmm) that eliminate runtime transposes and maximize throughput.
Microsoft's BitNet repository provides pretrained kernel parameters that optimize inference by pre-tuning weight matrix tiling sizes for specific GPU architectures. When you use these parameters, the system automatically configures exact tiling dimensions (BM, BK, and micro-tile bmm) during model conversion, ensuring the quantized weights align perfectly with the compiled kernels. This approach removes the need for manual profiling and delivers peak performance out-of-the-box for supported models like bitnet_b1_58-large and Llama-3 variants.
How Pretrained Kernel Parameters Work in BitNet
BitNet's inference performance depends on how weight matrices are tiled and quantized for GPU execution. The repository ships manually profiled configurations stored in preset_kernels/<model_name>/ directories, containing:
kernel_config_tl1.iniandkernel_config_tl2.ini: Configuration files mapping weight shapes (m, k) to optimal tile sizesbitnet-lut-kernels-tl1.handbitnet-lut-kernels-tl2.h: Pre-generated kernel headers compiled with these specific parameters
When enabled, setup_env.py copies these files into the include/ directory. During model conversion, utils/convert-hf-to-gguf-bitnet.py (lines 93-100) reads the configuration via ConfigParser and reshapes weights using preprocess_weights_tl1 to match the expected tiling, eliminating runtime transpose operations.
Enabling Pretuned Kernels During Setup
To activate pretrained kernel parameters, pass the --use-pretuned flag when running the environment setup script. The script automatically detects your model architecture and selects the appropriate TL-1 or TL-2 configuration.
python setup_env.py \
--hf-repo microsoft/bitnet-b1-58-large \
--quant-type tl1 \
--use-pretuned
This command performs three critical operations defined in setup_env.py:
- Detects the model architecture and locates the matching folder in
preset_kernels/ - Copies
bitnet-lut-kernels-tl1.htoinclude/bitnet-lut-kernels.h - Copies
kernel_config_tl1.initoinclude/kernel_config.ini
Converting Models with Kernel-Aware Weight Packing
After setup, convert your HuggingFace checkpoint using the pretuned configuration. The conversion script reads include/kernel_config.ini to determine exact reshaping parameters for each weight matrix.
python utils/convert-hf-to-gguf-bitnet.py \
--hf-repo microsoft/bitnet-b1-58-large \
--output-dir models/bitnet_b1_58-large
According to the source code in utils/convert-hf-to-gguf-bitnet.py (lines 93-99), the script iterates through kernel configuration sections to find entries where m and k match the weight dimensions, extracting BM, BK, and bmm values. These drive the preprocess_weights_tl1 pipeline that packs tensors into the exact format expected by the kernels.
Building and Running Optimized Inference
With kernel sources and configurations in place, setup_env.py invokes CMake to compile the inference binary using the pretuned tile sizes. The runtime (run_inference.py) then loads these pre-packed weights and calls the compiled kernels without additional overhead.
python run_inference.py \
--model-dir models/bitnet_b1_58-large \
--quant-type tl1 \
--prompt "What is the capital of France?"
Understanding Kernel Configuration File Structure
The .ini files in preset_kernels/<model>/ define tiling parameters for specific matrix shapes. Each section represents a kernel block:
[Kernels_0]
m = 1536
k = 4096
bm = 256
bk = 96
bmm = 32
- m and k: Input matrix dimensions (rows and columns)
- bm (BM): Block size in the M dimension
- bk (BK): Block size in the K dimension
- bmm: Micro-tile size for the matrix multiplication
These values are generated by utils/codegen_tl1.py and utils/codegen_tl2.py, which create the kernel headers based on the model's shape dictionary.
Loading Parameters Manually in Python
You can programmatically access pretuned parameters using the same ConfigParser logic found in the conversion utilities:
from configparser import ConfigParser
import os
def get_tile_params(m, k, config_path="include/kernel_config.ini"):
cfg = ConfigParser()
cfg.read(config_path)
for section in cfg.sections():
if int(cfg[section]["m"]) == m and int(cfg[section]["k"]) == k:
return {
"BM": int(cfg[section]["bm"]),
"BK": int(cfg[section]["bk"]),
"bmm": int(cfg[section]["bmm"]),
}
raise ValueError(f"No pretuned tile for shape ({m}, {k})")
# Retrieve parameters for a 1536×4096 weight matrix
params = get_tile_params(1536, 4096)
print(params) # Output: {'BM': 256, 'BK': 96, 'bmm': 32}
This matches the internal implementation used by preprocess_weights_tl1 to ensure weight tensors align with kernel expectations.
Summary
- Pretrained kernel parameters eliminate manual GPU kernel tuning by providing model-specific tile sizes (BM, BK, bmm) optimized for specific architectures.
- Activation method: Use the
--use-pretunedflag withsetup_env.pyto automatically copy configuration files frompreset_kernels/<model>/toinclude/. - Weight alignment: The
convert-hf-to-gguf-bitnet.pyscript readsinclude/kernel_config.iniviaConfigParser(lines 93-100) and reshapes weights usingpreprocess_weights_tl1to match kernel tiling. - File locations: Configurations reside in
preset_kernels/(e.g.,bitnet_b1_58-large/kernel_config_tl1.ini) and are generated byutils/codegen_tl1.pyandcodegen_tl2.py. - Performance impact: Proper tiling ensures the generated kernels consume tensors without runtime transposes, delivering maximum throughput during inference.
Frequently Asked Questions
What models support pretrained kernel parameters?
The preset_kernels/ directory contains optimized configurations for bitnet_b1_58-large and various Llama-3 variants. Each model folder includes both TL-1 and TL-2 quantization presets. Check the repository's preset_kernels/ directory for the current list of supported architectures.
How do BM, BK, and bmm parameters affect performance?
BM (block size M), BK (block size K), and bmm (micro-tile) determine how weight matrices are partitioned for GPU computation. These values are manually profiled to match specific GPU cache sizes and memory bandwidth characteristics. Incorrect values cause memory bank conflicts and reduced parallelism, while pretuned values maximize occupancy and minimize data movement.
Can I use pretuned kernels with custom model architectures?
No, pretuned kernels are model-specific because the tiling parameters depend on exact weight matrix dimensions. For custom architectures, you must either generate new kernel parameters using utils/codegen_tl1.py or codegen_tl2.py, or use the default auto-tuning path without --use-pretuned, though this may yield suboptimal performance compared to the manually profiled presets.
Where are the kernel source files located after setup?
When using --use-pretuned, setup_env.py copies bitnet-lut-kernels-tl1.h (or TL-2 variant) to include/bitnet-lut-kernels.h and kernel_config_tl1.ini to include/kernel_config.ini. The original files remain in preset_kernels/<model_name>/, preserving the source templates for reference.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →