# Fine-tuning GLM-5 for Domain-Specific Tasks: A Complete Guide to GLM-S Adaptation

> Master fine-tuning GLM-5 for your domain with GLM-S. Learn how parameter-efficient fine-tuning using LoRA adapters preserves the MoE architecture for optimal results.

- Repository: [Z.ai/GLM-5](https://github.com/zai-org/GLM-5)
- Tags: how-to-guide
- Published: 2026-06-19

---

**Fine-tuning GLM-5 for domain-specific tasks is most efficiently achieved through GLM-S, a lightweight variant that supports parameter-efficient fine-tuning (PEFT) with LoRA adapters while preserving the base model's sparse MoE architecture.**

The `zai-org/GLM-5` repository provides GLM-S as a streamlined entry point into the GLM-5 model family, specifically architected for rapid domain adaptation without requiring full model retraining. This guide demonstrates how to leverage the repository's modular skill ecosystem and HuggingFace integration to fine-tune GLM-S on specialized corpora using adapter-based methods.

## Understanding GLM-S Architecture for Fine-Tuning

GLM-S retains the core **General Language Model (GLM)** architecture found in the broader GLM-5 series, implementing a mixture-of-experts (MoE) transformer with **DeepSeek Sparse Attention (DSA)** and a **speculative decoding (MTP) layer** that enables long-context inference up to 1 million tokens.

### Sparse MoE with IndexShare

The model utilizes four sparse-attention layers that share a common indexer, reducing FLOPs by approximately 2.9× compared to dense attention while maintaining extensive context windows. In [`README.md`](https://github.com/zai-org/GLM-5/blob/main/README.md) (lines 20-26), this architecture is described as critical for efficient fine-tuning because the shared backbone can remain frozen while only lightweight adapter layers require training.

### Speculative Decoding and Reasoning Control

The MTP (Multi-Token Prediction) speculative decoding layer improves token acceptance length by up to 20% and exposes the `reasoning_effort` parameter, allowing you to toggle between `max` and `high` compute modes during inference. This is particularly valuable when evaluating fine-tuned checkpoints, as documented in [`README.md`](https://github.com/zai-org/GLM-5/blob/main/README.md) (line 80).

## Preparing Your Environment

Before fine-tuning, you must configure the GLM-S skill ecosystem and authentication credentials.

### Install the GLM-S Skill

The repository distributes GLM-S through a documentation-only master skill. Install it via Clawhub or clone directly from GitHub:

```bash
npx clawhub@latest install glm-s-skill

```

Alternatively, clone the repository containing the skill definitions from [`skills/glm-master-skill/SKILL.md`](https://github.com/zai-org/GLM-5/blob/main/skills/glm-master-skill/SKILL.md).

### Configure API Access

Set your `ZHIPU_API_KEY` environment variable, as required by the master skill documentation (lines 15-18):

```bash
export ZHIPU_API_KEY="your_key_here"

```

Install the required dependencies listed in [`requirements.txt`](https://github.com/zai-org/GLM-5/blob/main/requirements.txt):

```bash
pip install torch transformers peft datasets

```

## Fine-Tuning GLM-5 with LoRA Adapters

The recommended approach for domain-specific adaptation uses **Low-Rank Adaptation (LoRA)** to train small adapter matrices while keeping the base GLM-S parameters frozen. This preserves the model's general knowledge encoded in the sparse MoE layers while efficiently specializing output for your domain.

### Complete Fine-Tuning Script

Below is a minimal PyTorch implementation that loads the `zai-org/GLM-S` checkpoint, attaches LoRA adapters targeting the MoE attention layers, and trains on a custom dataset:

```python
import os
from datasets import load_dataset
from transformers import AutoTokenizer, AutoModelForCausalLM, TrainingArguments, Trainer
from peft import get_peft_model, LoraConfig

# -------------------------------------------------

# 1️⃣  Load tokenizer & base model (GLM‑S checkpoint)

# -------------------------------------------------

model_name = "zai-org/GLM-S"
tokenizer = AutoTokenizer.from_pretrained(model_name, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(model_name, trust_remote_code=True)

# -------------------------------------------------

# 2️⃣  Attach LoRA adapters (low‑rank fine‑tuning)

# -------------------------------------------------

lora_cfg = LoraConfig(
    r=8,                # rank

    lora_alpha=16,
    target_modules=["query_key_value"],  # typical for GLM MoE layers

    lora_dropout=0.1,
    bias="none",
)
model = get_peft_model(model, lora_cfg)

# -------------------------------------------------

# 3️⃣  Prepare a simple text dataset

# -------------------------------------------------

# Replace with your own domain data, e.g., a CSV with a "text" column

dataset = load_dataset("csv", data_files={
    "train": "data/domain_train.csv",
    "validation": "data/domain_val.csv"
})

def tokenize_fn(example):
    return tokenizer(example["text"], truncation=True, max_length=1024)

tokenized = dataset.map(tokenize_fn, batched=True, remove_columns=["text"])

# -------------------------------------------------

# 4️⃣  Define training arguments

# -------------------------------------------------

training_args = TrainingArguments(
    output_dir="./glm-s-finetuned",
    per_device_train_batch_size=2,
    per_device_eval_batch_size=2,
    num_train_epochs=3,
    learning_rate=5e-5,
    fp16=True,
    logging_steps=10,
    save_total_limit=2,
    evaluation_strategy="epoch",
)

# -------------------------------------------------

# 5️⃣  Initialise Trainer and start fine‑tuning

# -------------------------------------------------

trainer = Trainer(
    model=model,
    args=training_args,
    train_dataset=tokenized["train"],
    eval_dataset=tokenized["validation"],
)

trainer.train()
trainer.save_model("./glm-s-finetuned")

```

### Key Configuration Details

- **`target_modules=["query_key_value"]`**: This targets the attention projections within the sparse MoE layers, which is the standard configuration for GLM architecture fine-tuning.
- **`trust_remote_code=True`**: Required when loading `zai-org/GLM-S` to properly initialize the custom GLM model classes.
- **Checkpoint saving**: The script saves adapter weights to `./glm-s-finetuned`, which can be subsequently pushed to HuggingFace Hub for team collaboration.

## Domain Adaptation Best Practices

When adapting GLM-S to specialized domains such as medical, legal, or technical corpora:

1. **Keep the backbone frozen**: The sparse MoE layers in `zai-org/GLM-S` contain extensive general knowledge; freezing them prevents catastrophic forgetting.
2. **Use the [`example/ascend.md`](https://github.com/zai-org/GLM-5/blob/main/example/ascend.md) guide**: For hardware-specific deployments (e.g., Ascend NPU), consult this file in the repository's `example` folder to optimize training throughput.
3. **Leverage `reasoning_effort` during evaluation**: After fine-tuning, test your adapter with both `high` and `max` reasoning settings to determine the optimal compute-quality tradeoff for your domain.

## Summary

- **GLM-S** provides a lightweight pathway for fine-tuning GLM-5 on domain-specific tasks while retaining the full model's sparse attention capabilities.
- The architecture combines **IndexShare** sparse MoE layers with **MTP speculative decoding**, enabling efficient 1M-token context processing.
- **LoRA adapters** targeting `query_key_value` modules allow parameter-efficient fine-tuning without modifying the base model weights.
- Installation requires the **GLM-S skill** from [`skills/glm-master-skill/SKILL.md`](https://github.com/zai-org/GLM-5/blob/main/skills/glm-master-skill/SKILL.md) and a valid `ZHIPU_API_KEY` environment variable.
- The complete workflow involves loading the model with `trust_remote_code=True`, attaching PEFT adapters, and training via the HuggingFace Trainer API.

## Frequently Asked Questions

### What is the difference between GLM-5 and GLM-S?

GLM-S is a lightweight variant within the GLM-5 series designed specifically for rapid domain adaptation. While it maintains the same core MoE architecture with DeepSeek Sparse Attention as GLM-5 and GLM-5.1, it is optimized for scenarios where users need to fine-tune on domain-specific data without the computational overhead of the largest model variants.

### Do I need to unfreeze the MoE layers during fine-tuning?

No. According to the architecture documented in [`README.md`](https://github.com/zai-org/GLM-5/blob/main/README.md), you should keep the shared sparse backbone frozen and train only the adapter layers (LoRA or PEFT). This approach preserves the massive knowledge encoded in the base model while allowing efficient specialization through low-rank updates to the attention mechanisms.

### How do I handle long-context training data with GLM-S?

GLM-S supports contexts up to 1 million tokens through its IndexShare sparse attention mechanism. When preparing your dataset, use the tokenizer's `truncation=True` parameter with an appropriate `max_length` (e.g., 1024-4096 for most domain tasks, or higher if your hardware permits). The sparse attention layers in `zai-org/GLM-S` automatically handle the computational efficiency for longer sequences.

### Can I deploy fine-tuned GLM-S models on specialized hardware like Ascend NPU?

Yes. The repository includes [`example/ascend.md`](https://github.com/zai-org/GLM-5/blob/main/example/ascend.md), which provides specific guidance for deploying and running inference on Ascend hardware. The fine-tuning script remains compatible with standard PyTorch training loops, but the Ascend documentation covers hardware-specific optimizations for both training and inference phases.