Choosing Between FP8 and BF16 Precision Models for GLM-5 Deployment: A Complete Guide

Select BF16 for maximum accuracy and broad hardware compatibility, or choose FP8 for memory-constrained environments and specialized accelerators that support 8-bit inference.

The GLM-5 series (also referred to as GLM-S) from the zai-org/GLM-5 repository ships two precision variants for each model version, giving developers flexibility when deploying the 744-billion-parameter architecture. Both BF16 (Brain Float-16) and FP8 (Floating-Point 8) variants share identical model weights and architecture but differ in numerical representation during inference. Your choice between these formats depends on your hardware capabilities, memory constraints, and accuracy requirements.

GLM-5 Precision Variants Overview

The repository provides parallel downloads for each model size:

Model Precision Hardware Compatibility
GLM-5.2 BF16 General-purpose GPUs, CPUs, NPUs
GLM-5.2-FP8 FP8 Specialized ASICs (Huawei Ascend, custom NPUs)
GLM-5.1 BF16 Broad accelerator support
GLM-5.1-FP8 FP8 Memory-optimized deployments
GLM-5 BF16 Default production choice
GLM-5-FP8 FP8 High-throughput, low-RAM scenarios

Both variants maintain the full 744B parameter count and identical transformer architectures documented in README.md. The distinction lies solely in how activations and weights are quantized during the forward pass.

Why Choose BF16 (Brain Float-16)

BF16 is the default recommendation for most production deployments due to its balance of efficiency and numerical safety.

  • Numerical stability – BF16 retains the same 8-bit exponent as FP32, significantly reducing overflow and underflow risks. This is critical for long-context inference up to 1 million tokens, where accumulation errors can compound across layers.
  • Universal hardware support – Modern NVIDIA GPUs (A100, A40, H100), AMD MI250 accelerators, and Intel Xeon Scalable CPUs provide native BF16 matrix multiply units.
  • Minimal accuracy loss – Benchmarks show BF16 delivers less than 1% degradation relative to FP32 on standard reasoning tasks while providing approximately 2× speedup and 50% memory savings compared to full FP32.

Why Choose FP8 (Floating-Point 8)

FP8 becomes advantageous when maximizing throughput or operating under strict memory constraints.

  • Maximum memory compression – Storing values in single-byte precision reduces model weight footprint by 8× versus FP32 and 4× versus BF16. This enables fitting multiple model instances on a single device or serving higher concurrent request volumes.
  • Accelerator-specific throughput – Hardware like the Huawei Ascend NPU (documented in example/ascend.md) exposes FP8-specific kernels that can achieve greater than 3× the FLOPS of BF16 operations.
  • Acceptable accuracy retention – Empirical tests from the GLM-5 authors demonstrate that FP8 variants retain greater than 95% of BF16 scores on Terminal-Bench, SWE-Bench, and Vending-Bench benchmarks.

Decision Matrix for Production Deployment

Use this structured framework to select the appropriate precision for your infrastructure:

Decision Factor Choose BF16 Choose FP8
Hardware platform General-purpose GPUs (NVIDIA, AMD) or CPUs with BF16 support Proprietary NPUs or ASICs with optimized FP8 kernels
Memory constraints Standard VRAM configurations (80GB+ per instance) Low-RAM devices or multi-tenant serving requiring maximum model packing
Latency requirements Balanced throughput/accuracy Absolute lowest per-token latency with tolerable accuracy trade-off
Context length Long-context workloads (1M tokens) where numerical drift compounds Short to medium context tasks
Safety criticality Applications where small numerical errors could cascade Standard coding and reasoning tasks

Implementation Examples

The requirements.txt file specifies framework versions required for FP8 support, including vLLM ≥0.23.0 and SGLang ≥0.5.13.post1.

Loading with vLLM

from vllm import LLM, SamplingParams

# BF16 variant

llm_bf16 = LLM(model="zai-org/GLM-5.2", dtype="bfloat16")

# FP8 variant

llm_fp8 = LLM(model="zai-org/GLM-5.2-FP8", dtype="float8")

sampling = SamplingParams(temperature=0.7, max_new_tokens=256)
output = llm_fp8.generate(
    prompts=["Write a Python script to scrape a website"],
    sampling_params=sampling
)
print(output[0].text)

Loading with SGLang

import sglang as sg

# BF16 precision

model_bf16 = sg.Model("zai-org/GLM-5.2", precision="bf16")

# FP8 precision

model_fp8 = sg.Model("zai-org/GLM-5.2-FP8", precision="fp8")

response = model_fp8.generate("Summarize the key differences between FP8 and BF16.")
print(response)

Loading with Hugging Face Transformers

from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

def load_model(repo, precision):
    dtype = {"bf16": "bfloat16", "fp8": "float8"}[precision]
    tokenizer = AutoTokenizer.from_pretrained(repo)
    model = AutoModelForCausalLM.from_pretrained(
        repo, 
        torch_dtype=getattr(torch, dtype)
    )
    return tokenizer, model

tokenizer, model = load_model("zai-org/GLM-5-FP8", "fp8")
inputs = tokenizer("Explain why FP8 uses an 8-bit exponent", return_tensors="pt")
outputs = model.generate(**inputs, max_new_tokens=128)
print(tokenizer.decode(outputs[0]))

Key Source Files for Precision Selection

The zai-org/GLM-5 repository contains specific documentation for precision-related decisions:

  • README.md – Contains the official model table with download URLs for both BF16 and FP8 variants, plus high-level precision guidance.
  • example/ascend.md – Deployment guide specifically for Huawei Ascend NPU hardware, demonstrating FP8 inference optimizations.
  • requirements.txt – Lists minimum framework versions (vLLM ≥0.23.0, SGLang ≥0.5.13.post1) required for FP8 model loading.
  • skills/glm-master-skill/SKILL.md – Central navigation document that routes users to appropriate model repositories based on precision requirements.

Summary

  • BF16 provides the safest default for most GLM-5 deployments, offering broad hardware compatibility and minimal accuracy degradation for long-context and safety-critical workloads.
  • FP8 delivers maximum memory efficiency (4× smaller than BF16) and superior throughput on specialized ASICs like Huawei Ascend, retaining >95% of BF16 benchmark performance.
  • Both precision variants share the identical 744B parameter architecture and are available for GLM-5, GLM-5.1, and GLM-5.2 model families.
  • Framework support requires vLLM ≥0.23.0, SGLang ≥0.5.13.post1, or recent Hugging Face Transformers builds with torch.float8_e5m2 support.

Frequently Asked Questions

What is the parameter count for GLM-5 FP8 and BF16 models?

Both precision variants utilize the exact same 744 billion parameters. The FP8 and BF16 designations refer only to the inference-time numerical format for weights and activations, not to model size or architecture differences.

Which hardware platforms support FP8 inference for GLM-5?

FP8 inference requires specialized hardware such as Huawei Ascend NPUs (detailed in example/ascend.md) or proprietary AI accelerators with native 8-bit floating-point units. Standard NVIDIA and AMD GPUs should use the BF16 variants for optimal performance, as they lack efficient FP8 computation paths in most consumer and data center product lines.

How much accuracy degradation should I expect with FP8?

According to benchmarks published in the GLM-5 repository, FP8 models retain greater than 95% of BF16 performance on coding and reasoning benchmarks including Terminal-Bench, SWE-Bench, and Vending-Bench. For applications requiring maximum leaderboard scores or handling mathematical edge cases, BF16 remains the recommended choice.

Can I switch between BF16 and FP8 without changing my application code?

Yes, with minor configuration changes. The model architecture remains identical, so you only need to modify the dtype or precision parameter in your serving framework (vLLM, SGLang, or Transformers) and point to the appropriate Hugging Face or ModelScope repository (e.g., zai-org/GLM-5.2 for BF16 versus zai-org/GLM-5.2-FP8 for FP8). No changes to tokenization, prompt formatting, or generation logic are required.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →