What Models Are Officially Supported by BitNet: Complete List and Setup Guide
BitNet officially supports five model families—including bitnet_b1_58-large, bitnet_b1_58-3B, Llama3-8B-1.58-100B-tokens, Falcon3, and Falcon-E—each with specific kernel implementations for x86 and ARM architectures.
The microsoft/BitNet repository provides a curated list of verified models that work with the bitnet.cpp inference engine. These models are pre-configured with optimized kernels and can be downloaded automatically via the provided setup scripts.
Officially Supported Models
The repository maintains a definitive compatibility matrix in README.md (lines 86-154) that enumerates which models work with the CPU and GPU kernels. According to the source code, the following models are officially supported:
| Model | Parameters | x86 Support | ARM Support |
|---|---|---|---|
| bitnet_b1_58-large | 0.7 B | I2_S, TL2 | TL1, TL2 |
| bitnet_b1_58-3B | 3.3 B | TL2 | TL1 |
| Llama3-8B-1.58-100B-tokens | 8.0 B | I2_S, TL2 | TL1, TL2 |
| Falcon3 Family | 1 B – 10 B | I2_S, TL2 | TL1, TL2 |
| Falcon-E Family | 1 B – 3 B | I2_S, TL2 | TL1, TL2 |
The checkmarks indicate which quantization kernels are implemented and tested for each hardware platform. I2_S represents 2-bit integer weights with 16-bit accumulators, while TL1 and TL2 denote Ternary-Lookup kernels optimized for different memory-throughput trade-offs.
Kernel Support Matrix
Each supported model works with specific backend implementations in src/ggml-bitnet-*.cpp and gpu/bitnet_kernels/bitnet_kernels.cu. The kernel availability varies by architecture:
- I2_S: Available for bitnet_b1_58-large, Llama3-8B, Falcon3, and Falcon-E on x86; not available on ARM
- TL1: Available for bitnet_b1_58-3B and all models on ARM; not available on x86 for most models
- TL2: Universal support across all models on x86; available for bitnet_b1_58-large, Llama3-8B, Falcon3, and Falcon-E on ARM
The setup_env.py script (lines 14-18) validates these combinations when you specify a Hugging Face repository identifier.
How to Install Officially Supported Models
The setup_env.py utility restricts downloads to verified Hugging Face repositories. Valid --hf-repo values include:
1bitLLM/bitnet_b1_58-large1bitLLM/bitnet_b1_58-3BHF1BitLLM/Llama3-8B-1.58-100B-tokenstiiuae/Falcon3-1B-Instruct-1.58bit(and other Falcon3 variants)
To download an officially supported model:
# Install dependencies
conda create -n bitnet-cpp python=3.9
conda activate bitnet-cpp
pip install -r requirements.txt
# Download a verified model
python setup_env.py --hf-repo 1bitLLM/bitnet_b1_58-large
This command triggers automatic kernel generation via utils/codegen_tl1.py and utils/codegen_tl2.py, creating optimized lookup-table headers in preset_kernels/bitnet_b1_58-large/.
Running Inference on Supported Models
After installation, you can run inference using the pre-generated kernels. Each model ships with preset kernel configurations in preset_kernels/ directories containing tuned kernel_config_*.ini files.
CPU Inference
Use run_inference.py (defaulting to models/bitnet_b1_58-3B/ggml-model-i2_s.gguf per lines 45-48):
python run_inference.py \
-m models/bitnet_b1_58-large/ggml-model-i2_s.gguf \
-p "Explain the benefits of 1-bit LLMs."
To use TL1 or TL2 kernels instead, reference the corresponding GGUF file generated by the code generation utilities.
GPU Inference
For CUDA-accelerated inference, use run_inference_server.py (lines 49-55):
python run_inference_server.py \
-m models/bitnet_b1_58-large/ggml-model-tl2.gguf \
--port 8080
Query the server via HTTP:
curl -X POST http://localhost:8080/generate \
-d '{"prompt":"Summarize the BitNet paper in 2 sentences."}'
Model Architecture and Kernel Generation
Each officially supported model includes architecture-specific parameters that drive kernel generation. The utils/codegen_tl1.py and utils/codegen_tl2.py scripts automatically produce lookup-table headers (e.g., bitnet-lut-kernels-tl1.h) based on the model's hidden size and sequence length.
For testing without downloading full checkpoints, use the dummy model generator:
python utils/generate-dummy-bitnet-model.py \
models/bitnet_b1_58-large \
--outfile models/dummy-125m.tl1.gguf \
--outtype tl1 \
--model-size 125M
This creates a minimal GGUF file compatible with run_inference.py for pipeline verification.
Summary
- Five model families are officially supported: bitnet_b1_58-large (0.7B), bitnet_b1_58-3B (3.3B), Llama3-8B-1.58-100B-tokens (8B), Falcon3 (1B-10B), and Falcon-E (1B-3B).
- Kernel availability varies by architecture: x86 supports I2_S and TL2 for most models, while ARM primarily supports TL1 and TL2.
- Installation requires using
setup_env.pywith specific Hugging Face repository identifiers defined in the source code. - Inference scripts
run_inference.py(CPU) andrun_inference_server.py(GPU) automatically select appropriate kernels frompreset_kernels/directories.
Frequently Asked Questions
Can I use custom models with BitNet?
No, the setup_env.py script strictly validates --hf-repo arguments against the hardcoded dictionary in lines 14-18. While you can manually convert models using the quantization utilities, only the five officially supported families are guaranteed to work with the pre-generated kernels in preset_kernels/.
Which kernel should I use for maximum performance on x86?
TL2 (Ternary-Lookup-2) generally provides the best throughput on modern x86 vector units. According to the README compatibility table, TL2 is available for all supported models on x86 architectures, while I2_S is limited to specific models.
How do I verify my model is running the correct kernel?
Check the generated header files in preset_kernels/[model_name]/ after running setup_env.py. The presence of bitnet-lut-kernels-tl1.h or bitnet-lut-kernels-tl2.h confirms which lookup-table kernels were generated for your specific model architecture.
Are quantized Falcon models compatible with both CPU and GPU?
Yes, the Falcon3 and Falcon-E families support both CPU kernels (I2_S, TL2 on x86; TL1, TL2 on ARM) and CUDA kernels via run_inference_server.py. The same GGUF files work across both backends, though performance characteristics differ based on the selected kernel type.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →