Troubleshooting Slow Inference and OOM Errors with GLM-5: A Complete Optimization Guide

Resolve GLM-5 performance bottlenecks and out-of-memory crashes by loading the model in FP16 precision, caching the tokenizer singleton in utils.py, limiting sequence length to 512 tokens in config.yaml, and clearing the GPU cache after each inference cycle.

Slow inference and out-of-memory (OOM) errors in the zai-org/GLM-5 repository typically occur when running the GLM-S variant on consumer hardware with limited VRAM. The default configuration loads weights in full FP32 precision and instantiates the tokenizer repeatedly, causing unnecessary memory overhead. By modifying the loading logic in skills/glm-master-skill/glm_master.py and configuration values in config.yaml, you can reduce memory consumption by approximately 50% and prevent fragmentation-related crashes.

Optimize Model Precision and Loading

The GLM-S checkpoint defaults to FP32 (full precision), consuming approximately 2 GB of VRAM immediately upon loading. On GPUs with less than 8 GB of memory, this leaves insufficient room for activation maps during generation.

Load in Half-Precision

Configure the model loader in skills/glm-master-skill/glm_master.py to use torch.float16 or bfloat16 when supported:

model = AutoModelForCausalLM.from_pretrained(
    MODEL_ID,
    torch_dtype=torch.float16,  # Reduces VRAM by ~2x

    low_cpu_mem_usage=True    # Prevents loading full FP32 weights into system RAM

).to(DEVICE)

Setting low_cpu_mem_usage=True ensures that weights are not duplicated in system memory during the loading process, preventing CPU RAM exhaustion before the model even reaches the GPU.

Cache the Tokenizer

The default implementation may re-instantiate the tokenizer on every request, adding significant CPU overhead. Implement a singleton pattern in skills/glm-master-skill/utils.py to cache the tokenizer at module import:


# utils.py

tokenizer = AutoTokenizer.from_pretrained(MODEL_ID)

def get_tokenizer():
    return tokenizer

This eliminates redundant loading latency and reduces memory fragmentation from repeated object creation.

Manage Sequence Length and Batch Configuration

Large prompt lengths (≥1024 tokens) create massive activation maps even with batch size 1. The total memory requirement scales with batch_size × sequence_length, quickly exhausting available VRAM on long inputs.

Truncate Prompts in Configuration

Set a conservative maximum length in skills/glm-master-skill/config.yaml to prevent OOM during preprocessing:

MAX_TOKENS = 512  # Adjust based on your device's VRAM headroom

input_ids = tokenizer.encode(prompt, truncation=True, max_length=MAX_TOKENS)

For interactive chat applications, 512 tokens provides sufficient context while keeping activation memory under control.

Minimize CPU-GPU Transfer Overhead

Every inference call moves input tensors from host memory to device memory. Under concurrent load, PCIe bandwidth becomes a bottleneck, and gradient graph construction adds unnecessary overhead.

Use No-Grad Context and Asynchronous Transfers

Modify the generate_reply function in skills/glm-master-skill/glm_master.py to disable gradient computation and enable non-blocking memory transfers:

with torch.no_grad():
    input_tensor = torch.tensor(input_ids).to(device, non_blocking=True)
    output = model.generate(input_tensor, **gen_kwargs)

The torch.no_grad() context prevents PyTorch from building computation graphs for backpropagation, reducing memory overhead by 15-20% during inference. The non_blocking=True flag allows the CPU to continue processing while the GPU transfers data, improving throughput for batched requests.

Prevent Memory Fragmentation

Long-running services suffer from CUDA memory fragmentation, where available memory exists but is not contiguous, triggering OOM errors despite sufficient total VRAM.

Clear Cache After Each Request

Add torch.cuda.empty_cache() to your inference cleanup routine to defragment memory without restarting the service:

def safe_inference(prompt: str):
    try:
        reply = generate(prompt)
        return reply
    finally:
        torch.cuda.empty_cache()  # Lightweight fragmentation fix

For production deployments, monitor fragmentation levels with nvidia-smi and restart the service periodically if memory utilization creeps upward over time.

System Environment Tuning

Operating system configuration significantly impacts inference latency.

Disable Swap and Verify CUDA Support

Linux swap files cause severe latency spikes when the system attempts to page GPU-accessible memory. Ensure swap is disabled for production inference:

sudo swapoff -a
cat /proc/meminfo | grep Swap

Verify your PyTorch installation recognizes CUDA to prevent silent CPU fallback:

python -c "import torch; print(torch.__version__); print(torch.cuda.is_available())"

Use the latest CUDA driver compatible with your GPU architecture to ensure access to optimized kernels.

Complete Implementation Examples

Example 1: Optimized Model Loading with FP16

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

DEVICE = torch.device("cuda" if torch.cuda.is_available() else "cpu")
MODEL_ID = "zai-org/glm-5-s"

# Cache tokenizer and model at module import

tokenizer = AutoTokenizer.from_pretrained(MODEL_ID)
model = AutoModelForCausalLM.from_pretrained(
    MODEL_ID,
    torch_dtype=torch.float16,
    low_cpu_mem_usage=True
).to(DEVICE)

def generate(prompt: str, max_new_tokens: int = 128):
    input_ids = tokenizer.encode(
        prompt, 
        return_tensors="pt", 
        truncation=True, 
        max_length=512
    ).to(DEVICE)
    
    with torch.no_grad():
        output_ids = model.generate(
            input_ids,
            max_new_tokens=max_new_tokens,
            temperature=0.7,
            do_sample=True,
        )
    return tokenizer.decode(output_ids[0], skip_special_tokens=True)

Example 2: Safe Inference with Memory Cleanup

def safe_inference(prompt: str):
    """Wrapper to ensure memory is freed after each call."""
    try:
        return generate(prompt)
    finally:
        torch.cuda.empty_cache()

Example 3: Adaptive Token Limit Based on Available VRAM

def adaptive_generate(prompt: str):
    """Dynamically adjust generation length based on free GPU memory."""
    # Heuristic: 1 token ≈ 2KB in FP16 activation space

    free_mem_mb = torch.cuda.mem_get_info()[0] / 1024**2
    safe_max_tokens = int(free_mem_mb / 2)  # Reserve headroom for kv-cache

    
    return generate(prompt, max_new_tokens=min(256, safe_max_tokens))

Summary

  • Load in FP16: Configure torch_dtype=torch.float16 and low_cpu_mem_usage=True in glm_master.py to halve VRAM consumption.
  • Cache Tokenizer: Implement a singleton pattern in utils.py to eliminate redundant instantiation overhead.
  • Limit Sequence Length: Set MAX_TOKENS=512 in config.yaml and use truncation=True during encoding.
  • Optimize Transfers: Wrap inference in torch.no_grad() and use non_blocking=True for CPU-to-GPU tensor movement.
  • Clear Fragmentation: Call torch.cuda.empty_cache() after each generation cycle to prevent memory fragmentation crashes.
  • System Tuning: Disable Linux swap and verify CUDA driver compatibility to ensure optimal kernel performance.

Frequently Asked Questions

Why does GLM-5 run out of memory on a 6GB GPU despite the model being only 2GB?

The GLM-S checkpoint consumes approximately 2GB in FP32 precision, but generation requires additional memory for activation maps and the key-value cache, which scale with sequence length. At 1024 tokens, these activations can exceed 4GB. Load the model in FP16 via torch_dtype=torch.float16 in skills/glm-master-skill/glm_master.py to reduce the base footprint to ~1GB, leaving sufficient room for longer contexts.

How do I fix slow inference speed in the GLM-5 skill implementation?

Slow inference typically results from redundant tokenizer loading and unnecessary gradient computation. Cache the tokenizer as a module-level singleton in skills/glm-master-skill/utils.py and wrap the model.generate() call in a torch.no_grad() context. These changes eliminate CPU overhead and reduce memory allocation time by 15-20%.

What causes CUDA out-of-memory errors after several hours of running GLM-5?

Memory fragmentation occurs when PyTorch's CUDA allocator cannot find contiguous blocks of memory despite having sufficient total VRAM available. Call torch.cuda.empty_cache() after each inference request in your generate_reply function to defragment the memory pool. For services running continuously, implement a periodic restart mechanism if monitoring shows creeping memory usage.

Should I use FP16 or BF16 for GLM-5 inference?

Use FP16 (torch.float16) for broad hardware compatibility, as it is supported on all NVIDIA GPUs from Pascal onwards. Use BF16 (torch.bfloat16) only if running on Ampere or newer architectures (A100, RTX 30-series) where it provides better numerical stability with the same memory efficiency. Both options reduce the model footprint from ~2GB to ~1GB according to the GLM-5 source configuration.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →