BitNet b1.58 vs 4-Bit Activation LLMs: How the 8-Bit bI.S8 Variant Works

BitNet b1.58 with 8-bit activations (bI.S8) trades modest memory overhead for significantly faster inference by replacing per-token weight unpacking with lookup-table kernels, delivering 1.2–1.5× higher throughput on long prompts compared to 4-bit activation variants.

The microsoft/BitNet repository implements a mixed-precision architecture that maintains the ultra-compact 1-bit weight format while upgrading activations from 4-bit to 8-bit integers. This shift fundamentally changes the performance-memory trade-off, favoring latency-sensitive inference workloads over minimal activation footprints.

BitNet b1.58 Activation Formats: From 4-Bit to 8-Bit

BitNet b1.58 models use binary weights paired with varying activation precisions. The bI.S8 variant specifically employs 1-bit weights with 8-bit signed integer (int8) activations, contrasting with other implementations in the family:

  • TL1/TL2: 1-bit weights with 4-bit integer (int4) activations using parallel weight-unpack kernels
  • I2_S: 1-bit weights with 2-bit signed integer (int2) activations using native GEMM/GEMV kernels
  • bI.S8: 1-bit weights with 8-bit integer (int8) activations using lookup-table (LUT) kernels

The bI.S8 format targets scenarios where activation unpacking dominates runtime, particularly during long-prompt inference.

Lookup-Table Kernel Architecture

The bI.S8 variant achieves speed through Lookup-Table (LUT) kernels implemented in src/ggml-bitnet-lut.cpp and the multiply-add (MAD) logic in src/ggml-bitnet-mad.cpp. These kernels pre-compute all 256 possible products of a binary weight and 8-bit activation value.

Because the weight matrix unpacks once and reuses across many int8 activation vectors, the kernel eliminates repeated branchy logic. The LUT remains resident in registers (256 entries), avoiding memory bottlenecks during operations defined by the QK_I2_S macro within the MAD kernel source.

Performance and Memory Trade-offs

Compared to 4-bit activation variants, bI.S8 exhibits distinct operational characteristics:

Memory footprint: bI.S8 consumes roughly 2× more activation RAM than TL1/TL2 int4 variants, though still significantly less than full-precision FP16 models.

Throughput: On CPUs with wide SIMD units (AVX-512, ARM SVE) and GPUs with strong integer pipelines, bI.S8 achieves 1.2–1.5× higher token-per-second rates for long prompts. The elimination of per-element activation unpacking reduces cycle count per token.

Use case alignment: While 4-bit activation variants target edge devices with tight memory constraints, bI.S8 optimizes for in-situ inference where latency dominates over the modest extra activation memory.

Building and Converting to bI.S8

Build the BitNet Runtime

git clone --recursive https://github.com/microsoft/BitNet.git
cd BitNet

conda create -n bitnet-cpp python=3.9 -y
conda activate bitnet-cpp
pip install -r requirements.txt

mkdir -p build && cd build
cmake .. -DCMAKE_BUILD_TYPE=Release
make -j$(nproc)

Convert Models to 8-Bit Activations

The utils/convert-helper-bitnet.py wrapper exposes the --quant-type i8 flag to generate bI.S8 checkpoints:

huggingface-cli download microsoft/BitNet-b1.58-2B-4T-gguf --local-dir models/BitNet-2B

python utils/convert-helper-bitnet.py \
    models/BitNet-2B/ggml-model-f32.gguf \
    --quant-type i8

This preserves the 1-bit weight file while quantizing the activation buffer to int8, as implemented in the conversion helper logic.

Run Inference with bI.S8

python run_inference.py \
    -m models/BitNet-2B/ggml-model-i8.gguf \
    -p "Explain quantization-aware training." \
    -t 8

Switch between precision levels by changing the model file:

  • ggml-model-i2_s.gguf for 2-bit activations
  • ggml-model-tl1.gguf for 4-bit activations
  • ggml-model-i8.gguf for 8-bit activations (bI.S8)

Summary

  • BitNet b1.58 bI.S8 uses 1-bit weights with 8-bit int8 activations, unlike 4-bit activation variants that use TL1/TL2 formats.
  • LUT kernels in src/ggml-bitnet-lut.cpp eliminate per-token unpacking overhead, improving long-prompt throughput by 1.2–1.5× on modern SIMD hardware.
  • Memory cost is approximately 2× higher than 4-bit variants, making bI.S8 ideal for latency-critical server inference rather than memory-constrained edge deployment.
  • Configuration uses --quant-type i8 via utils/convert-helper-bitnet.py to enable the 8-bit activation path.

Frequently Asked Questions

What is the difference between BitNet b1.58 and BitNet bI.S8?

BitNet b1.58 refers to the model family using compressed binary weights. BitNet bI.S8 is a specific variant within this family that pairs 1-bit weights with 8-bit integer activations, as opposed to other variants using 4-bit or 2-bit activations for different performance-memory trade-offs.

Why would I choose 8-bit activations over 4-bit for a 1-bit LLM?

Choose 8-bit activations (bI.S8) when inference latency matters more than activation memory usage. The lookup-table kernels in src/ggml-bitnet-mad.cpp eliminate unpacking overhead, yielding 1.2–1.5× faster token generation on long prompts compared to 4-bit variants that require parallel weight-unpack operations per token.

How do I convert an existing BitNet model to use bI.S8 activations?

Use the conversion wrapper at utils/convert-helper-bitnet.py with the --quant-type i8 flag. This tool processes the original FP32 weights and generates a ggml-model-i8.gguf file that maintains 1-bit weight compression while quantizing the activation pathway to 8-bit integers.

Does bI.S8 work on both CPU and GPU?

Yes. The LUT kernels are implemented for both CPU (AVX-512, ARM SVE) and GPU architectures in src/ggml-bitnet-lut.cpp and src/ggml-bitnet-mad.cpp. Performance gains are most pronounced on hardware with wide SIMD units or strong integer pipelines.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →