# How to Use Unsloth for Fine-Tuning GLM-5.2: A Complete Guide

> Learn to fine-tune GLM-5.2 with Unsloth. This guide details a memory-efficient LoRA workflow for consumer GPUs, preserving IndexShare sparsity and MTP speculative decoding.

- Repository: [Z.ai/GLM-5](https://github.com/zai-org/GLM-5)
- Tags: how-to-guide
- Published: 2026-06-21

---

**Unsloth provides a memory-efficient LoRA workflow that integrates seamlessly with GLM-5.2, allowing you to fine-tune the model on consumer GPUs by wrapping the base checkpoint with lightweight adapters while preserving the underlying IndexShare sparsity and MTP speculative decoding layers.**

The zai-org/GLM-5 repository officially supports Unsloth for parameter-efficient fine-tuning of the GLM-5.2 checkpoint. By leveraging Unsloth's optimized LoRA implementation, you can fine-tune this large language model without modifying the base weights or requiring excessive GPU memory.

## Prerequisites and Installation

Before starting, ensure you have the minimum required versions. According to the GLM-5 repository documentation, you need **Unsloth** ≥ `0.1.47-beta`, **Transformers** ≥ `0.5.12`, PyTorch, and Accelerate.

Install the dependencies with pip:

```bash
pip install "unsloth[accelerate]" transformers>=0.5.12 torch accelerate

```

These versions ensure compatibility with GLM-5.2's architecture and Unsloth's LoRA injection mechanisms.

## Loading the GLM-5.2 Base Model

Load the base model from Hugging Face or ModelScope using the standard Transformers API. The GLM-5 team hosts the checkpoints under the `zai-org/GLM-5.2` namespace.

```python
from transformers import AutoTokenizer, AutoModelForCausalLM

tokenizer = AutoTokenizer.from_pretrained("zai-org/GLM-5.2")

base_model = AutoModelForCausalLM.from_pretrained(
    "zai-org/GLM-5.2",
    device_map="auto",          # Automatic device placement across GPUs

    torch_dtype="auto",         # Automatically selects fp16 or bf16

)

```

The `device_map="auto"` parameter handles multi-GPU distribution, while `torch_dtype="auto"` selects the appropriate precision based on your hardware capabilities.

## Applying Unsloth LoRA Configuration

Unsloth wraps the base model using `LoRAConfig` and `prepare_model_for_lora`. This approach injects low-rank adapters into specific attention projection matrices without altering the original weights.

```python
from unsloth import LoRAConfig, prepare_model_for_lora

lora_cfg = LoRAConfig(
    r=64,                        # Rank of the adaptation matrices

    lora_alpha=16,               # Scaling parameter for the LoRA updates

    target_modules=["q_proj", "v_proj"],  # GLM-5 attention heads to adapt

)

model = prepare_model_for_lora(base_model, lora_cfg)

```

Targeting **q_proj** and **v_proj** specifically allows the adapter to modify the query and value projections while leaving the IndexShare sparsity layers and MTP speculative decoding components untouched.

## Preparing Your Fine-Tuning Dataset

Convert your instruction-following data into the Hugging Face `datasets` format. The following example uses an Alpaca-style JSON file:

```python
from datasets import load_dataset

data = load_dataset("json", data_files="my_dataset.json")

def tokenize_fn(example):
    return tokenizer(
        example["instruction"],
        truncation=True,
        max_length=1024,
    )

tokenized = data.map(tokenize_fn, batched=True)

```

Ensure your dataset contains the fields expected by your training loop, typically `instruction` and `input` for standard fine-tuning tasks.

## Configuring the Training Loop

Unsloth reuses the standard Hugging Face `Trainer` but automatically handles gradient accumulation and optimizer sharding for LoRA. Configure the training arguments with mixed precision enabled:

```python
from transformers import Trainer, TrainingArguments

training_args = TrainingArguments(
    output_dir="./glm5_finetuned",
    per_device_train_batch_size=4,
    gradient_accumulation_steps=8,
    num_train_epochs=3,
    learning_rate=2e-4,
    fp16=True,                  # Enable mixed-precision training

    logging_steps=10,
    save_steps=500,
    eval_strategy="no",
)

trainer = Trainer(
    model=model,
    args=training_args,
    train_dataset=tokenized["train"],
)

trainer.train()

```

The `gradient_accumulation_steps=8` setting effectively simulates a batch size of 32 while keeping memory usage low.

## Saving and Loading LoRA Adapters

After training, persist only the adapter parameters rather than the full model weights. This keeps storage requirements minimal and allows for efficient model switching.

```python
model.save_pretrained("./glm5_finetuned_lora")
tokenizer.save_pretrained("./glm5_finetuned_lora")

```

To resume training or deploy the model later, load the base checkpoint and merge the saved LoRA weights using the same configuration.

## Running Inference with Fine-Tuned Weights

Load your fine-tuned adapter for text generation. The GLM-5.2 model supports the `reasoning_effort` parameter, documented in the repository's README, to control chain-of-thought compute intensity.

```python
from transformers import pipeline

generator = pipeline(
    "text-generation",
    model="./glm5_finetuned_lora",
    tokenizer=tokenizer,
    max_new_tokens=256,
    temperature=0.7,
)

output = generator(
    "Explain the quantum advantage of GLM-5.2:",
    do_sample=True,
    reasoning_effort="high",    # Optional: increases reasoning compute

)

print(output[0]["generated_text"])

```

Setting `reasoning_effort="high"` instructs the model to spend additional computation on complex reasoning tasks, as implemented in the zai-org/GLM-5 source code.

## Summary

- **Installation**: Use Unsloth ≥ `0.1.47-beta` and Transformers ≥ `0.5.12` for full GLM-5.2 compatibility.
- **Model Loading**: Load checkpoints via `zai-org/GLM-5.2` with automatic device mapping and dtype detection.
- **LoRA Configuration**: Wrap models using `LoRAConfig` targeting `q_proj` and `v_proj` to preserve sparsity layers.
- **Training**: Utilize standard `Trainer` with gradient accumulation; only adapter weights update during training.
- **Inference**: Support for `reasoning_effort` parameter persists through fine-tuning for controlled reasoning depth.

## Frequently Asked Questions

### What hardware is required to fine-tune GLM-5.2 with Unsloth?

Unsloth's LoRA implementation reduces memory requirements significantly, allowing fine-tuning on consumer GPUs with as little as 16GB VRAM depending on batch size and sequence length. The `device_map="auto"` and `gradient_accumulation_steps` parameters help distribute workloads across multiple GPUs if available.

### How does Unsloth preserve GLM-5.2's architectural optimizations?

Unsloth injects adapters only into the attention projection matrices (`q_proj`, `v_proj`), leaving the **IndexShare** sparsity modules and **MTP** (Multi-Token Prediction) speculative decoding layers untouched. This architectural preservation ensures the model retains its original inference speed and memory efficiency while gaining task-specific adaptations.

### Can I use the reasoning_effort parameter with fine-tuned models?

Yes, the `reasoning_effort` parameter remains available during inference after fine-tuning. This parameter, referenced in the GLM-5 repository's README at lines 78-80, controls how much computational effort the model expends on chain-of-thought reasoning. Set it to `"high"`, `"medium"`, or `"low"` when calling the generation pipeline.

### Where can I find the official model checkpoints for GLM-5.2?

The official download links are listed in the **Download Model** table in the repository's README at lines 61-68. Both Hugging Face and ModelScope mirrors are available under the `zai-org/GLM-5.2` namespace, ensuring reliable access regardless of your geographic location.