BitNet Lookup Table (LUT) Methodology: Accelerating 2-Bit Quantized Inference

BitNet’s lookup table (LUT) methodology pre-computes 8-bit weight contributions for every possible 2-bit pattern, allowing inference kernels to replace costly bit-extraction and multiplication with single-instruction vectorized table lookups using ARM NEON or x86 SIMD intrinsics.

The microsoft/BitNet repository implements ultra-fast quantized matrix multiplication by treating 2-bit weights as indices into pre-computed lookup tables rather than performing runtime dequantization. This LUT methodology transforms a bandwidth-bound bit manipulation task into a compute-bound register operation, delivering significant speedups on both ARM and x86 architectures.

How BitNet Builds and Uses Lookup Tables

The LUT methodology operates through four distinct stages that convert floating-point weights into cache-friendly lookup structures optimized for vectorized hardware instructions.

Stage 1: Quantization and Scale Computation

The process begins with per-tensor quantization where FP32 weights are compressed to 2-bit integers (values in {-2, -1, 0, 1}) and associated with a scaling factor.

In preset_kernels/bitnet_b1_58-large/bitnet-lut-kernels-tl1.h, the per_tensor_quant function (lines 24-31) computes the scale as lut_scales = 127 / max(|w|), normalizing the quantized range to maximize precision while maintaining the 2-bit constraint. This scale is stored alongside the LUT for later dequantization during inference.

Stage 2: Packing int2 Values into int8 Bytes

To enable efficient vectorized access, the 2-bit values must be densely packed. The compress_int2_to_int8 function in gpu/pack_weight.py (lines 46-53) packs four 2-bit weights into a single byte, creating a compressed layout that matches the indexing scheme used by the lookup table.

This packing reduces memory bandwidth by 4× and aligns the data structure with the 8-bit vector registers used in subsequent SIMD operations.

Stage 3: Constructing the Tile-Based Lookup Table

BitNet processes matrices in 16×32 tiles (where wmma_n = 16 and wmma_k = 32). For each tile, the system builds a mapping from local coordinates to shared memory layouts.

The B_global_16x32_to_shared_load_16x32_layout function in gpu/pack_weight.py (lines 5-14) generates the coordinate mapping, while lut_ctor in bitnet-lut-kernels-tl1.h (lines 95-115) converts this mapping into NEON-friendly 8-bit vectors. The resulting lookup table stores pre-computed row and column offsets for every possible 2-bit pattern combination, fitting approximately 256 bytes per tile into L1 cache.

Stage 4: Vectorized Inference Using LUT Lookup

During inference, the generated kernel (qgemm_lut_*) loads activation tiles and performs the matrix multiplication via table lookup rather than arithmetic.

The kernel extracts low and high nibbles from each activation byte, then uses vqtbl1q_s8 (ARM NEON) or equivalent x86 instructions to retrieve the pre-computed 8-bit contributions from the LUT. The tbl_impl_1536_4096 function in bitnet-lut-kernels-tl1.h (lines 184-210) implements this core logic, accumulating results into 32-bit integers before applying lut_scales to convert back to FP16/FP32 via ggml_qgemm_lut in include/ggml-bitnet.h (lines 38-44).

Why the LUT Methodology Improves Performance

The lookup table approach provides three critical advantages over naive 2-bit dequantization:

  • Eliminated Per-Bit Computation: Instead of extracting 2-bit values and performing FP16 multiplication for every weight, the kernel executes a single vector-table-lookup instruction per byte (covering four weights).
  • Register-Local Data Access: The LUT resides in CPU registers (int8x16_t vec_lut[…]) and is accessed via SIMD intrinsics like vqtbl1q_s8, enabling single-instruction, data-parallel operations without memory bottlenecks.
  • Cache Efficiency: With a footprint of roughly 256 bytes per 16×32 tile, the LUT remains resident in L1 cache throughout the computation, eliminating the bandwidth pressure typical of quantized matrix multiplication.

Implementation Examples

Generating Packed Weights and LUT in Python

The following Python workflow converts FP32 weights into the packed LUT format required by BitNet kernels:

import numpy as np
from bitnet.gpu.pack_weight import permutate_weight_fastest, compress_int2_to_int8, interleave_weight_int8

# Assume `weight` is a (N, K) float32 matrix from a transformer layer

int2_weight = (weight * 127 / np.max(np.abs(weight)))  # naive int2 quantisation

int2_weight = np.clip(np.rint(int2_weight), -2, 1).astype(np.int8)  # values in {-2,-1,0,1}

# 1. Permute weight into 16×32 tiles (GPU-friendly layout)

perm = permutate_weight_fastest(int2_weight)

# 2. Pack 2-bit values into int8-bytes

packed = compress_int2_to_int8(perm)

# 3. Interleave bytes for the final LUT layout

lut = interleave_weight_int8(packed, nbits=2)   # shape (N, K/4) int8

# `lut` can now be fed to ggml_bitnet_transform_tensor() via the C API

Running Inference with the LUT in C

The high-level C API abstracts the LUT complexity while leveraging the optimized kernels:

#include "ggml-bitnet.h"

int main(void) {
    // Load a BitNet model (weights already contain the LUT)
    struct ggml_context * ctx = ggml_init(/*...*/);

    // `tensor` is a quantised weight tensor that contains the LUT & scales
    struct ggml_tensor * weight = ggml_get_tensor(ctx, "model.layers.0.attn.q_proj.weight");

    // Initialise BitNet runtime (sets up LUT buffers, etc.)
    ggml_bitnet_init();

    // Typical matmul: output = weight @ input
    struct ggml_tensor * input  = ggml_new_tensor_2d(ctx, GGML_TYPE_F32, dim_in, batch);
    struct ggml_tensor * output = ggml_mul_mat(ctx, weight, input);

    // The kernel behind `ggml_mul_mat` will dispatch `ggml_qgemm_lut`
    // which reads the pre-computed LUT and performs the fast lookup
    // multiplication described above.

    ggml_bitnet_free();   // clean up LUT buffers
    ggml_free(ctx);
    return 0;
}

Key Source Files and Functions

Understanding the LUT methodology requires familiarity with these specific components in the microsoft/BitNet repository:

Summary

  • BitNet’s LUT methodology pre-computes 8-bit weight contributions for 2-bit quantized values, storing them in tile-specific lookup tables.
  • The implementation processes matrices in 16×32 tiles, using vqtbl1q_s8 (ARM) or AVX-512 equivalents to perform single-instruction vectorized lookups.
  • Key functions include per_tensor_quant for scaling, compress_int2_to_int8 for packing, and ggml_qgemm_lut for inference dispatch.
  • The approach eliminates per-bit arithmetic, fitting ~256-byte LUTs into L1 cache for high-throughput quantized matrix multiplication.

Frequently Asked Questions

What is the size of each lookup table in BitNet?

Each lookup table corresponds to a 16×32 tile of weights and occupies approximately 256 bytes. This compact size ensures the LUT remains resident in L1 cache during the matrix multiplication kernel execution, eliminating memory bandwidth bottlenecks.

How does BitNet handle the 2-bit to 8-bit conversion during inference?

BitNet avoids runtime conversion by using vectorized table lookups. The kernel loads packed int8 bytes containing four 2-bit weights each, extracts the high and low nibbles, and uses the vqtbl1q_s8 intrinsic (or x86 equivalent) to look up pre-computed 8-bit contributions from the LUT, accumulating results in 32-bit integers before final scaling.

Can the LUT methodology work on x86 processors without ARM NEON?

Yes. BitNet provides architecture-specific implementations in bitnet-lut-kernels-tl2.h for x86 platforms, utilizing AVX-2 and AVX-512 instructions such as vpmovqb to achieve the same single-instruction lookup performance as the ARM NEON implementation in bitnet-lut-kernels-tl1.h.

Where are the LUT scales stored in a BitNet model?

The per-tensor scaling factors (lut_scales) are computed during quantization by per_tensor_quant in bitnet-lut-kernels-tl1.h and stored alongside the packed weight tensors. During inference, ggml_qgemm_lut (declared in include/ggml-bitnet.h) applies these scales to convert accumulated integer results back to floating-point precision.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →