How to Use Unsloth for Fine-Tuning GLM-5.2: A Complete Guide

Unsloth provides a memory-efficient LoRA workflow that integrates seamlessly with GLM-5.2, allowing you to fine-tune the model on consumer GPUs by wrapping the base checkpoint with lightweight adapters while preserving the underlying IndexShare sparsity and MTP speculative decoding layers.

The zai-org/GLM-5 repository officially supports Unsloth for parameter-efficient fine-tuning of the GLM-5.2 checkpoint. By leveraging Unsloth's optimized LoRA implementation, you can fine-tune this large language model without modifying the base weights or requiring excessive GPU memory.

Prerequisites and Installation

Before starting, ensure you have the minimum required versions. According to the GLM-5 repository documentation, you need Unsloth ≥ 0.1.47-beta, Transformers ≥ 0.5.12, PyTorch, and Accelerate.

Install the dependencies with pip:

pip install "unsloth[accelerate]" transformers>=0.5.12 torch accelerate

These versions ensure compatibility with GLM-5.2's architecture and Unsloth's LoRA injection mechanisms.

Loading the GLM-5.2 Base Model

Load the base model from Hugging Face or ModelScope using the standard Transformers API. The GLM-5 team hosts the checkpoints under the zai-org/GLM-5.2 namespace.

from transformers import AutoTokenizer, AutoModelForCausalLM

tokenizer = AutoTokenizer.from_pretrained("zai-org/GLM-5.2")

base_model = AutoModelForCausalLM.from_pretrained(
    "zai-org/GLM-5.2",
    device_map="auto",          # Automatic device placement across GPUs

    torch_dtype="auto",         # Automatically selects fp16 or bf16

)

The device_map="auto" parameter handles multi-GPU distribution, while torch_dtype="auto" selects the appropriate precision based on your hardware capabilities.

Applying Unsloth LoRA Configuration

Unsloth wraps the base model using LoRAConfig and prepare_model_for_lora. This approach injects low-rank adapters into specific attention projection matrices without altering the original weights.

from unsloth import LoRAConfig, prepare_model_for_lora

lora_cfg = LoRAConfig(
    r=64,                        # Rank of the adaptation matrices

    lora_alpha=16,               # Scaling parameter for the LoRA updates

    target_modules=["q_proj", "v_proj"],  # GLM-5 attention heads to adapt

)

model = prepare_model_for_lora(base_model, lora_cfg)

Targeting q_proj and v_proj specifically allows the adapter to modify the query and value projections while leaving the IndexShare sparsity layers and MTP speculative decoding components untouched.

Preparing Your Fine-Tuning Dataset

Convert your instruction-following data into the Hugging Face datasets format. The following example uses an Alpaca-style JSON file:

from datasets import load_dataset

data = load_dataset("json", data_files="my_dataset.json")

def tokenize_fn(example):
    return tokenizer(
        example["instruction"],
        truncation=True,
        max_length=1024,
    )

tokenized = data.map(tokenize_fn, batched=True)

Ensure your dataset contains the fields expected by your training loop, typically instruction and input for standard fine-tuning tasks.

Configuring the Training Loop

Unsloth reuses the standard Hugging Face Trainer but automatically handles gradient accumulation and optimizer sharding for LoRA. Configure the training arguments with mixed precision enabled:

from transformers import Trainer, TrainingArguments

training_args = TrainingArguments(
    output_dir="./glm5_finetuned",
    per_device_train_batch_size=4,
    gradient_accumulation_steps=8,
    num_train_epochs=3,
    learning_rate=2e-4,
    fp16=True,                  # Enable mixed-precision training

    logging_steps=10,
    save_steps=500,
    eval_strategy="no",
)

trainer = Trainer(
    model=model,
    args=training_args,
    train_dataset=tokenized["train"],
)

trainer.train()

The gradient_accumulation_steps=8 setting effectively simulates a batch size of 32 while keeping memory usage low.

Saving and Loading LoRA Adapters

After training, persist only the adapter parameters rather than the full model weights. This keeps storage requirements minimal and allows for efficient model switching.

model.save_pretrained("./glm5_finetuned_lora")
tokenizer.save_pretrained("./glm5_finetuned_lora")

To resume training or deploy the model later, load the base checkpoint and merge the saved LoRA weights using the same configuration.

Running Inference with Fine-Tuned Weights

Load your fine-tuned adapter for text generation. The GLM-5.2 model supports the reasoning_effort parameter, documented in the repository's README, to control chain-of-thought compute intensity.

from transformers import pipeline

generator = pipeline(
    "text-generation",
    model="./glm5_finetuned_lora",
    tokenizer=tokenizer,
    max_new_tokens=256,
    temperature=0.7,
)

output = generator(
    "Explain the quantum advantage of GLM-5.2:",
    do_sample=True,
    reasoning_effort="high",    # Optional: increases reasoning compute

)

print(output[0]["generated_text"])

Setting reasoning_effort="high" instructs the model to spend additional computation on complex reasoning tasks, as implemented in the zai-org/GLM-5 source code.

Summary

  • Installation: Use Unsloth ≥ 0.1.47-beta and Transformers ≥ 0.5.12 for full GLM-5.2 compatibility.
  • Model Loading: Load checkpoints via zai-org/GLM-5.2 with automatic device mapping and dtype detection.
  • LoRA Configuration: Wrap models using LoRAConfig targeting q_proj and v_proj to preserve sparsity layers.
  • Training: Utilize standard Trainer with gradient accumulation; only adapter weights update during training.
  • Inference: Support for reasoning_effort parameter persists through fine-tuning for controlled reasoning depth.

Frequently Asked Questions

What hardware is required to fine-tune GLM-5.2 with Unsloth?

Unsloth's LoRA implementation reduces memory requirements significantly, allowing fine-tuning on consumer GPUs with as little as 16GB VRAM depending on batch size and sequence length. The device_map="auto" and gradient_accumulation_steps parameters help distribute workloads across multiple GPUs if available.

How does Unsloth preserve GLM-5.2's architectural optimizations?

Unsloth injects adapters only into the attention projection matrices (q_proj, v_proj), leaving the IndexShare sparsity modules and MTP (Multi-Token Prediction) speculative decoding layers untouched. This architectural preservation ensures the model retains its original inference speed and memory efficiency while gaining task-specific adaptations.

Can I use the reasoning_effort parameter with fine-tuned models?

Yes, the reasoning_effort parameter remains available during inference after fine-tuning. This parameter, referenced in the GLM-5 repository's README at lines 78-80, controls how much computational effort the model expends on chain-of-thought reasoning. Set it to "high", "medium", or "low" when calling the generation pipeline.

Where can I find the official model checkpoints for GLM-5.2?

The official download links are listed in the Download Model table in the repository's README at lines 61-68. Both Hugging Face and ModelScope mirrors are available under the zai-org/GLM-5.2 namespace, ensuring reliable access regardless of your geographic location.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →