# BitNet Lookup Table (LUT) Methodology: Accelerating 2-Bit Quantized Inference

> Explore BitNet's lookup table LUT methodology. Discover how pre-computing 8-bit weight contributions for 2-bit patterns accelerates inference with vectorized table lookups.

- Repository: [Microsoft/BitNet](https://github.com/microsoft/BitNet)
- Tags: deep-dive
- Published: 2026-03-13

---

**BitNet’s lookup table (LUT) methodology pre-computes 8-bit weight contributions for every possible 2-bit pattern, allowing inference kernels to replace costly bit-extraction and multiplication with single-instruction vectorized table lookups using ARM NEON or x86 SIMD intrinsics.**

The microsoft/BitNet repository implements ultra-fast quantized matrix multiplication by treating 2-bit weights as indices into pre-computed lookup tables rather than performing runtime dequantization. This **LUT methodology** transforms a bandwidth-bound bit manipulation task into a compute-bound register operation, delivering significant speedups on both ARM and x86 architectures.

## How BitNet Builds and Uses Lookup Tables

The LUT methodology operates through four distinct stages that convert floating-point weights into cache-friendly lookup structures optimized for vectorized hardware instructions.

### Stage 1: Quantization and Scale Computation

The process begins with **per-tensor quantization** where FP32 weights are compressed to 2-bit integers (values in {-2, -1, 0, 1}) and associated with a scaling factor.

In [`preset_kernels/bitnet_b1_58-large/bitnet-lut-kernels-tl1.h`](https://github.com/microsoft/BitNet/blob/main/preset_kernels/bitnet_b1_58-large/bitnet-lut-kernels-tl1.h), the `per_tensor_quant` function (lines 24-31) computes the scale as `lut_scales = 127 / max(|w|)`, normalizing the quantized range to maximize precision while maintaining the 2-bit constraint. This scale is stored alongside the LUT for later dequantization during inference.

### Stage 2: Packing int2 Values into int8 Bytes

To enable efficient vectorized access, the 2-bit values must be densely packed. The `compress_int2_to_int8` function in [`gpu/pack_weight.py`](https://github.com/microsoft/BitNet/blob/main/gpu/pack_weight.py) (lines 46-53) packs four 2-bit weights into a single byte, creating a compressed layout that matches the indexing scheme used by the lookup table.

This packing reduces memory bandwidth by 4× and aligns the data structure with the 8-bit vector registers used in subsequent SIMD operations.

### Stage 3: Constructing the Tile-Based Lookup Table

BitNet processes matrices in **16×32 tiles** (where `wmma_n = 16` and `wmma_k = 32`). For each tile, the system builds a mapping from local coordinates to shared memory layouts.

The `B_global_16x32_to_shared_load_16x32_layout` function in [`gpu/pack_weight.py`](https://github.com/microsoft/BitNet/blob/main/gpu/pack_weight.py) (lines 5-14) generates the coordinate mapping, while `lut_ctor` in [`bitnet-lut-kernels-tl1.h`](https://github.com/microsoft/BitNet/blob/main/bitnet-lut-kernels-tl1.h) (lines 95-115) converts this mapping into NEON-friendly 8-bit vectors. The resulting lookup table stores pre-computed row and column offsets for every possible 2-bit pattern combination, fitting approximately 256 bytes per tile into L1 cache.

### Stage 4: Vectorized Inference Using LUT Lookup

During inference, the generated kernel (`qgemm_lut_*`) loads activation tiles and performs the matrix multiplication via table lookup rather than arithmetic.

The kernel extracts low and high nibbles from each activation byte, then uses `vqtbl1q_s8` (ARM NEON) or equivalent x86 instructions to retrieve the pre-computed 8-bit contributions from the LUT. The `tbl_impl_1536_4096` function in [`bitnet-lut-kernels-tl1.h`](https://github.com/microsoft/BitNet/blob/main/bitnet-lut-kernels-tl1.h) (lines 184-210) implements this core logic, accumulating results into 32-bit integers before applying `lut_scales` to convert back to FP16/FP32 via `ggml_qgemm_lut` in [`include/ggml-bitnet.h`](https://github.com/microsoft/BitNet/blob/main/include/ggml-bitnet.h) (lines 38-44).

## Why the LUT Methodology Improves Performance

The lookup table approach provides three critical advantages over naive 2-bit dequantization:

- **Eliminated Per-Bit Computation**: Instead of extracting 2-bit values and performing FP16 multiplication for every weight, the kernel executes a single vector-table-lookup instruction per byte (covering four weights).
- **Register-Local Data Access**: The LUT resides in CPU registers (`int8x16_t vec_lut[…]`) and is accessed via SIMD intrinsics like `vqtbl1q_s8`, enabling single-instruction, data-parallel operations without memory bottlenecks.
- **Cache Efficiency**: With a footprint of roughly 256 bytes per 16×32 tile, the LUT remains resident in L1 cache throughout the computation, eliminating the bandwidth pressure typical of quantized matrix multiplication.

## Implementation Examples

### Generating Packed Weights and LUT in Python

The following Python workflow converts FP32 weights into the packed LUT format required by BitNet kernels:

```python
import numpy as np
from bitnet.gpu.pack_weight import permutate_weight_fastest, compress_int2_to_int8, interleave_weight_int8

# Assume `weight` is a (N, K) float32 matrix from a transformer layer

int2_weight = (weight * 127 / np.max(np.abs(weight)))  # naive int2 quantisation

int2_weight = np.clip(np.rint(int2_weight), -2, 1).astype(np.int8)  # values in {-2,-1,0,1}

# 1. Permute weight into 16×32 tiles (GPU-friendly layout)

perm = permutate_weight_fastest(int2_weight)

# 2. Pack 2-bit values into int8-bytes

packed = compress_int2_to_int8(perm)

# 3. Interleave bytes for the final LUT layout

lut = interleave_weight_int8(packed, nbits=2)   # shape (N, K/4) int8

# `lut` can now be fed to ggml_bitnet_transform_tensor() via the C API

```

### Running Inference with the LUT in C

The high-level C API abstracts the LUT complexity while leveraging the optimized kernels:

```c
#include "ggml-bitnet.h"

int main(void) {
    // Load a BitNet model (weights already contain the LUT)
    struct ggml_context * ctx = ggml_init(/*...*/);

    // `tensor` is a quantised weight tensor that contains the LUT & scales
    struct ggml_tensor * weight = ggml_get_tensor(ctx, "model.layers.0.attn.q_proj.weight");

    // Initialise BitNet runtime (sets up LUT buffers, etc.)
    ggml_bitnet_init();

    // Typical matmul: output = weight @ input
    struct ggml_tensor * input  = ggml_new_tensor_2d(ctx, GGML_TYPE_F32, dim_in, batch);
    struct ggml_tensor * output = ggml_mul_mat(ctx, weight, input);

    // The kernel behind `ggml_mul_mat` will dispatch `ggml_qgemm_lut`
    // which reads the pre-computed LUT and performs the fast lookup
    // multiplication described above.

    ggml_bitnet_free();   // clean up LUT buffers
    ggml_free(ctx);
    return 0;
}

```

## Key Source Files and Functions

Understanding the LUT methodology requires familiarity with these specific components in the microsoft/BitNet repository:

- **[`gpu/pack_weight.py`](https://github.com/microsoft/BitNet/blob/main/gpu/pack_weight.py)**: Contains `compress_int2_to_int8` and `B_global_16x32_to_shared_load_16x32_layout`, which generate the permutation mapping and pack int2 values into the LUT layout.
- **[`preset_kernels/bitnet_b1_58-large/bitnet-lut-kernels-tl1.h`](https://github.com/microsoft/BitNet/blob/main/preset_kernels/bitnet_b1_58-large/bitnet-lut-kernels-tl1.h)**: Implements ARM NEON kernels including `per_tensor_quant`, `lut_ctor`, and `tbl_impl_*` functions for LUT construction and matrix multiplication.
- **[`preset_kernels/bitnet_b1_58-large/bitnet-lut-kernels-tl2.h`](https://github.com/microsoft/BitNet/blob/main/preset_kernels/bitnet_b1_58-large/bitnet-lut-kernels-tl2.h)**: Provides the x86 AVX-2/AVX-512 equivalent kernels using the same LUT methodology.
- **[`include/ggml-bitnet.h`](https://github.com/microsoft/BitNet/blob/main/include/ggml-bitnet.h)**: Declares the public API functions `ggml_bitnet_init`, `ggml_qgemm_lut`, and `ggml_preprocessor` for LUT-based inference.
- **[`src/ggml-bitnet-lut.cpp`](https://github.com/microsoft/BitNet/blob/main/src/ggml-bitnet-lut.cpp)**: Contains the platform-specific implementation and dispatch logic for `ggml_qgemm_lut` across ARM (TL1) and x86 (TL2) backends.
- **[`utils/codegen_tl1.py`](https://github.com/microsoft/BitNet/blob/main/utils/codegen_tl1.py) and [`utils/codegen_tl2.py`](https://github.com/microsoft/BitNet/blob/main/utils/codegen_tl2.py)**: Python generators that emit optimized LUT kernels from high-level tile descriptions.

## Summary

- BitNet’s **LUT methodology** pre-computes 8-bit weight contributions for 2-bit quantized values, storing them in tile-specific lookup tables.
- The implementation processes matrices in **16×32 tiles**, using `vqtbl1q_s8` (ARM) or AVX-512 equivalents to perform single-instruction vectorized lookups.
- Key functions include `per_tensor_quant` for scaling, `compress_int2_to_int8` for packing, and `ggml_qgemm_lut` for inference dispatch.
- The approach eliminates per-bit arithmetic, fitting ~256-byte LUTs into L1 cache for high-throughput quantized matrix multiplication.

## Frequently Asked Questions

### What is the size of each lookup table in BitNet?

Each lookup table corresponds to a **16×32 tile** of weights and occupies approximately **256 bytes**. This compact size ensures the LUT remains resident in L1 cache during the matrix multiplication kernel execution, eliminating memory bandwidth bottlenecks.

### How does BitNet handle the 2-bit to 8-bit conversion during inference?

BitNet avoids runtime conversion by using **vectorized table lookups**. The kernel loads packed int8 bytes containing four 2-bit weights each, extracts the high and low nibbles, and uses the `vqtbl1q_s8` intrinsic (or x86 equivalent) to look up pre-computed 8-bit contributions from the LUT, accumulating results in 32-bit integers before final scaling.

### Can the LUT methodology work on x86 processors without ARM NEON?

Yes. BitNet provides architecture-specific implementations in [`bitnet-lut-kernels-tl2.h`](https://github.com/microsoft/BitNet/blob/main/bitnet-lut-kernels-tl2.h) for x86 platforms, utilizing **AVX-2 and AVX-512** instructions such as `vpmovqb` to achieve the same single-instruction lookup performance as the ARM NEON implementation in [`bitnet-lut-kernels-tl1.h`](https://github.com/microsoft/BitNet/blob/main/bitnet-lut-kernels-tl1.h).

### Where are the LUT scales stored in a BitNet model?

The per-tensor scaling factors (`lut_scales`) are computed during quantization by `per_tensor_quant` in [`bitnet-lut-kernels-tl1.h`](https://github.com/microsoft/BitNet/blob/main/bitnet-lut-kernels-tl1.h) and stored alongside the packed weight tensors. During inference, `ggml_qgemm_lut` (declared in [`include/ggml-bitnet.h`](https://github.com/microsoft/BitNet/blob/main/include/ggml-bitnet.h)) applies these scales to convert accumulated integer results back to floating-point precision.