Choosing Between FP8 and BF16 Precision Models for GLM-5 Deployment: A Complete Guide
Select BF16 for maximum accuracy and broad hardware compatibility, or choose FP8 for memory-constrained environments and specialized accelerators that support 8-bit inference.
The GLM-5 series (also referred to as GLM-S) from the zai-org/GLM-5 repository ships two precision variants for each model version, giving developers flexibility when deploying the 744-billion-parameter architecture. Both BF16 (Brain Float-16) and FP8 (Floating-Point 8) variants share identical model weights and architecture but differ in numerical representation during inference. Your choice between these formats depends on your hardware capabilities, memory constraints, and accuracy requirements.
GLM-5 Precision Variants Overview
The repository provides parallel downloads for each model size:
| Model | Precision | Hardware Compatibility |
|---|---|---|
| GLM-5.2 | BF16 | General-purpose GPUs, CPUs, NPUs |
| GLM-5.2-FP8 | FP8 | Specialized ASICs (Huawei Ascend, custom NPUs) |
| GLM-5.1 | BF16 | Broad accelerator support |
| GLM-5.1-FP8 | FP8 | Memory-optimized deployments |
| GLM-5 | BF16 | Default production choice |
| GLM-5-FP8 | FP8 | High-throughput, low-RAM scenarios |
Both variants maintain the full 744B parameter count and identical transformer architectures documented in README.md. The distinction lies solely in how activations and weights are quantized during the forward pass.
Why Choose BF16 (Brain Float-16)
BF16 is the default recommendation for most production deployments due to its balance of efficiency and numerical safety.
- Numerical stability – BF16 retains the same 8-bit exponent as FP32, significantly reducing overflow and underflow risks. This is critical for long-context inference up to 1 million tokens, where accumulation errors can compound across layers.
- Universal hardware support – Modern NVIDIA GPUs (A100, A40, H100), AMD MI250 accelerators, and Intel Xeon Scalable CPUs provide native BF16 matrix multiply units.
- Minimal accuracy loss – Benchmarks show BF16 delivers less than 1% degradation relative to FP32 on standard reasoning tasks while providing approximately 2× speedup and 50% memory savings compared to full FP32.
Why Choose FP8 (Floating-Point 8)
FP8 becomes advantageous when maximizing throughput or operating under strict memory constraints.
- Maximum memory compression – Storing values in single-byte precision reduces model weight footprint by 8× versus FP32 and 4× versus BF16. This enables fitting multiple model instances on a single device or serving higher concurrent request volumes.
- Accelerator-specific throughput – Hardware like the Huawei Ascend NPU (documented in
example/ascend.md) exposes FP8-specific kernels that can achieve greater than 3× the FLOPS of BF16 operations. - Acceptable accuracy retention – Empirical tests from the GLM-5 authors demonstrate that FP8 variants retain greater than 95% of BF16 scores on Terminal-Bench, SWE-Bench, and Vending-Bench benchmarks.
Decision Matrix for Production Deployment
Use this structured framework to select the appropriate precision for your infrastructure:
| Decision Factor | Choose BF16 | Choose FP8 |
|---|---|---|
| Hardware platform | General-purpose GPUs (NVIDIA, AMD) or CPUs with BF16 support | Proprietary NPUs or ASICs with optimized FP8 kernels |
| Memory constraints | Standard VRAM configurations (80GB+ per instance) | Low-RAM devices or multi-tenant serving requiring maximum model packing |
| Latency requirements | Balanced throughput/accuracy | Absolute lowest per-token latency with tolerable accuracy trade-off |
| Context length | Long-context workloads (1M tokens) where numerical drift compounds | Short to medium context tasks |
| Safety criticality | Applications where small numerical errors could cascade | Standard coding and reasoning tasks |
Implementation Examples
The requirements.txt file specifies framework versions required for FP8 support, including vLLM ≥0.23.0 and SGLang ≥0.5.13.post1.
Loading with vLLM
from vllm import LLM, SamplingParams
# BF16 variant
llm_bf16 = LLM(model="zai-org/GLM-5.2", dtype="bfloat16")
# FP8 variant
llm_fp8 = LLM(model="zai-org/GLM-5.2-FP8", dtype="float8")
sampling = SamplingParams(temperature=0.7, max_new_tokens=256)
output = llm_fp8.generate(
prompts=["Write a Python script to scrape a website"],
sampling_params=sampling
)
print(output[0].text)
Loading with SGLang
import sglang as sg
# BF16 precision
model_bf16 = sg.Model("zai-org/GLM-5.2", precision="bf16")
# FP8 precision
model_fp8 = sg.Model("zai-org/GLM-5.2-FP8", precision="fp8")
response = model_fp8.generate("Summarize the key differences between FP8 and BF16.")
print(response)
Loading with Hugging Face Transformers
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
def load_model(repo, precision):
dtype = {"bf16": "bfloat16", "fp8": "float8"}[precision]
tokenizer = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(
repo,
torch_dtype=getattr(torch, dtype)
)
return tokenizer, model
tokenizer, model = load_model("zai-org/GLM-5-FP8", "fp8")
inputs = tokenizer("Explain why FP8 uses an 8-bit exponent", return_tensors="pt")
outputs = model.generate(**inputs, max_new_tokens=128)
print(tokenizer.decode(outputs[0]))
Key Source Files for Precision Selection
The zai-org/GLM-5 repository contains specific documentation for precision-related decisions:
README.md– Contains the official model table with download URLs for both BF16 and FP8 variants, plus high-level precision guidance.example/ascend.md– Deployment guide specifically for Huawei Ascend NPU hardware, demonstrating FP8 inference optimizations.requirements.txt– Lists minimum framework versions (vLLM ≥0.23.0, SGLang ≥0.5.13.post1) required for FP8 model loading.skills/glm-master-skill/SKILL.md– Central navigation document that routes users to appropriate model repositories based on precision requirements.
Summary
- BF16 provides the safest default for most GLM-5 deployments, offering broad hardware compatibility and minimal accuracy degradation for long-context and safety-critical workloads.
- FP8 delivers maximum memory efficiency (4× smaller than BF16) and superior throughput on specialized ASICs like Huawei Ascend, retaining >95% of BF16 benchmark performance.
- Both precision variants share the identical 744B parameter architecture and are available for GLM-5, GLM-5.1, and GLM-5.2 model families.
- Framework support requires vLLM ≥0.23.0, SGLang ≥0.5.13.post1, or recent Hugging Face Transformers builds with
torch.float8_e5m2support.
Frequently Asked Questions
What is the parameter count for GLM-5 FP8 and BF16 models?
Both precision variants utilize the exact same 744 billion parameters. The FP8 and BF16 designations refer only to the inference-time numerical format for weights and activations, not to model size or architecture differences.
Which hardware platforms support FP8 inference for GLM-5?
FP8 inference requires specialized hardware such as Huawei Ascend NPUs (detailed in example/ascend.md) or proprietary AI accelerators with native 8-bit floating-point units. Standard NVIDIA and AMD GPUs should use the BF16 variants for optimal performance, as they lack efficient FP8 computation paths in most consumer and data center product lines.
How much accuracy degradation should I expect with FP8?
According to benchmarks published in the GLM-5 repository, FP8 models retain greater than 95% of BF16 performance on coding and reasoning benchmarks including Terminal-Bench, SWE-Bench, and Vending-Bench. For applications requiring maximum leaderboard scores or handling mathematical edge cases, BF16 remains the recommended choice.
Can I switch between BF16 and FP8 without changing my application code?
Yes, with minor configuration changes. The model architecture remains identical, so you only need to modify the dtype or precision parameter in your serving framework (vLLM, SGLang, or Transformers) and point to the appropriate Hugging Face or ModelScope repository (e.g., zai-org/GLM-5.2 for BF16 versus zai-org/GLM-5.2-FP8 for FP8). No changes to tokenization, prompt formatting, or generation logic are required.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →