How to Use Unsloth for Fine-Tuning GLM-5.2: A Complete Guide
Unsloth provides a memory-efficient LoRA workflow that integrates seamlessly with GLM-5.2, allowing you to fine-tune the model on consumer GPUs by wrapping the base checkpoint with lightweight adapters while preserving the underlying IndexShare sparsity and MTP speculative decoding layers.
The zai-org/GLM-5 repository officially supports Unsloth for parameter-efficient fine-tuning of the GLM-5.2 checkpoint. By leveraging Unsloth's optimized LoRA implementation, you can fine-tune this large language model without modifying the base weights or requiring excessive GPU memory.
Prerequisites and Installation
Before starting, ensure you have the minimum required versions. According to the GLM-5 repository documentation, you need Unsloth ≥ 0.1.47-beta, Transformers ≥ 0.5.12, PyTorch, and Accelerate.
Install the dependencies with pip:
pip install "unsloth[accelerate]" transformers>=0.5.12 torch accelerate
These versions ensure compatibility with GLM-5.2's architecture and Unsloth's LoRA injection mechanisms.
Loading the GLM-5.2 Base Model
Load the base model from Hugging Face or ModelScope using the standard Transformers API. The GLM-5 team hosts the checkpoints under the zai-org/GLM-5.2 namespace.
from transformers import AutoTokenizer, AutoModelForCausalLM
tokenizer = AutoTokenizer.from_pretrained("zai-org/GLM-5.2")
base_model = AutoModelForCausalLM.from_pretrained(
"zai-org/GLM-5.2",
device_map="auto", # Automatic device placement across GPUs
torch_dtype="auto", # Automatically selects fp16 or bf16
)
The device_map="auto" parameter handles multi-GPU distribution, while torch_dtype="auto" selects the appropriate precision based on your hardware capabilities.
Applying Unsloth LoRA Configuration
Unsloth wraps the base model using LoRAConfig and prepare_model_for_lora. This approach injects low-rank adapters into specific attention projection matrices without altering the original weights.
from unsloth import LoRAConfig, prepare_model_for_lora
lora_cfg = LoRAConfig(
r=64, # Rank of the adaptation matrices
lora_alpha=16, # Scaling parameter for the LoRA updates
target_modules=["q_proj", "v_proj"], # GLM-5 attention heads to adapt
)
model = prepare_model_for_lora(base_model, lora_cfg)
Targeting q_proj and v_proj specifically allows the adapter to modify the query and value projections while leaving the IndexShare sparsity layers and MTP speculative decoding components untouched.
Preparing Your Fine-Tuning Dataset
Convert your instruction-following data into the Hugging Face datasets format. The following example uses an Alpaca-style JSON file:
from datasets import load_dataset
data = load_dataset("json", data_files="my_dataset.json")
def tokenize_fn(example):
return tokenizer(
example["instruction"],
truncation=True,
max_length=1024,
)
tokenized = data.map(tokenize_fn, batched=True)
Ensure your dataset contains the fields expected by your training loop, typically instruction and input for standard fine-tuning tasks.
Configuring the Training Loop
Unsloth reuses the standard Hugging Face Trainer but automatically handles gradient accumulation and optimizer sharding for LoRA. Configure the training arguments with mixed precision enabled:
from transformers import Trainer, TrainingArguments
training_args = TrainingArguments(
output_dir="./glm5_finetuned",
per_device_train_batch_size=4,
gradient_accumulation_steps=8,
num_train_epochs=3,
learning_rate=2e-4,
fp16=True, # Enable mixed-precision training
logging_steps=10,
save_steps=500,
eval_strategy="no",
)
trainer = Trainer(
model=model,
args=training_args,
train_dataset=tokenized["train"],
)
trainer.train()
The gradient_accumulation_steps=8 setting effectively simulates a batch size of 32 while keeping memory usage low.
Saving and Loading LoRA Adapters
After training, persist only the adapter parameters rather than the full model weights. This keeps storage requirements minimal and allows for efficient model switching.
model.save_pretrained("./glm5_finetuned_lora")
tokenizer.save_pretrained("./glm5_finetuned_lora")
To resume training or deploy the model later, load the base checkpoint and merge the saved LoRA weights using the same configuration.
Running Inference with Fine-Tuned Weights
Load your fine-tuned adapter for text generation. The GLM-5.2 model supports the reasoning_effort parameter, documented in the repository's README, to control chain-of-thought compute intensity.
from transformers import pipeline
generator = pipeline(
"text-generation",
model="./glm5_finetuned_lora",
tokenizer=tokenizer,
max_new_tokens=256,
temperature=0.7,
)
output = generator(
"Explain the quantum advantage of GLM-5.2:",
do_sample=True,
reasoning_effort="high", # Optional: increases reasoning compute
)
print(output[0]["generated_text"])
Setting reasoning_effort="high" instructs the model to spend additional computation on complex reasoning tasks, as implemented in the zai-org/GLM-5 source code.
Summary
- Installation: Use Unsloth ≥
0.1.47-betaand Transformers ≥0.5.12for full GLM-5.2 compatibility. - Model Loading: Load checkpoints via
zai-org/GLM-5.2with automatic device mapping and dtype detection. - LoRA Configuration: Wrap models using
LoRAConfigtargetingq_projandv_projto preserve sparsity layers. - Training: Utilize standard
Trainerwith gradient accumulation; only adapter weights update during training. - Inference: Support for
reasoning_effortparameter persists through fine-tuning for controlled reasoning depth.
Frequently Asked Questions
What hardware is required to fine-tune GLM-5.2 with Unsloth?
Unsloth's LoRA implementation reduces memory requirements significantly, allowing fine-tuning on consumer GPUs with as little as 16GB VRAM depending on batch size and sequence length. The device_map="auto" and gradient_accumulation_steps parameters help distribute workloads across multiple GPUs if available.
How does Unsloth preserve GLM-5.2's architectural optimizations?
Unsloth injects adapters only into the attention projection matrices (q_proj, v_proj), leaving the IndexShare sparsity modules and MTP (Multi-Token Prediction) speculative decoding layers untouched. This architectural preservation ensures the model retains its original inference speed and memory efficiency while gaining task-specific adaptations.
Can I use the reasoning_effort parameter with fine-tuned models?
Yes, the reasoning_effort parameter remains available during inference after fine-tuning. This parameter, referenced in the GLM-5 repository's README at lines 78-80, controls how much computational effort the model expends on chain-of-thought reasoning. Set it to "high", "medium", or "low" when calling the generation pipeline.
Where can I find the official model checkpoints for GLM-5.2?
The official download links are listed in the Download Model table in the repository's README at lines 61-68. Both Hugging Face and ModelScope mirrors are available under the zai-org/GLM-5.2 namespace, ensuring reliable access regardless of your geographic location.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →