# Controlling Reasoning Effort in GLM-5 API: max vs high

> Master GLM-5 API reasoning effort control. Choose between max for speed or high for deeper Chain-of-Thought analysis to optimize your results.

- Repository: [Z.ai/GLM-5](https://github.com/zai-org/GLM-5)
- Tags: how-to-guide
- Published: 2026-06-19

---

**The GLM-5 API provides a `reasoning_effort` parameter that lets you trade latency for deeper reasoning by setting `"max"` (default, faster) or `"high"` (more thorough Chain-of-Thought).**

The `zai-org/GLM-5` repository exposes a request-level control for reasoning depth that directly impacts how the GLM-5, GLM-5.2, and GLM-5.1 models allocate internal compute resources. By adjusting the `reasoning_effort` parameter, developers can optimize for either low-latency responses or more accurate, deeply-reasoned outputs without modifying model weights or deployment infrastructure.

## Understanding the reasoning_effort Parameter

The `reasoning_effort` parameter accepts string values that determine the budget for internal Chain-of-Thought (CoT) processing. According to the source documentation in [`README.md`](https://github.com/zai-org/GLM-5/blob/main/README.md), this parameter is parsed at the API gateway and passed downstream to the inference engine.

### max (Default)

**`"max"`** is the default setting that activates a standard reasoning budget optimized for speed-to-answer. When this value is set, the model uses a bounded thinking token that stops after a preset number of internal steps, guaranteeing consistent latency for general inference tasks. This level is ideal for benchmark reproductions, production workloads, and real-time applications where response time is critical.

### high

**`"high"`** instructs the model to allocate a larger reasoning budget, keeping the "thinking" token active for significantly more steps. In this mode, the transformer decoder performs more extensive internal Chain-of-Thought steps and searches the knowledge store more thoroughly. This produces more accurate or creative answers but increases latency due to additional compute allocation. Use this for complex coding problems, multi-step planning tasks, or when the default output lacks sufficient depth.

### Disabling Thinking Entirely

Set **`enable_thinking=false`** to completely disable the model's internal CoT reasoning. This returns the fastest possible answer by bypassing the thinking token entirely, making it suitable for real-time UI feedback or extremely low-resource environments where any reasoning overhead is unacceptable.

## How Reasoning Effort Works Internally

When the request arrives at the serving layer, the JSON payload is parsed and the `reasoning_effort` value triggers specific mechanisms in the inference engine. As implemented in `zai-org/GLM-5`, the `high` setting lifts the thinking budget flag, which influences two primary subsystems:

1. **Sparse Attention Budget** – The Dynamic Sparse Attention (DSA) module allocates additional compute to the "thinking" token, allowing it to attend to more context windows and maintain longer coherent reasoning chains.

2. **Iterative Tool Calls** – In agentic modes, the model may schedule extra tool calls (such as code execution or search operations) before finalizing its response. The `"high"` level permits more such iterations, while the default `"max"` level enforces stricter limits to maintain performance guarantees.

The parameter validation occurs in the API gateway logic before being passed to the model inference process, ensuring no additional parameters need adjustment when switching between modes.

## Implementation Examples

### Python with Requests

Use the Z.ai API endpoint to control reasoning effort programmatically:

```python
import requests
import json

url = "https://api.z.ai/v1/chat/completions"
headers = {
    "Authorization": f"Bearer {YOUR_ZHIPU_API_KEY}",
    "Content-Type": "application/json"
}

payload = {
    "model": "glm-5.2",  # or glm-5.1 / glm-5

    "messages": [
        {"role": "user", "content": "Explain the difference between recursion and iteration."}
    ],
    "reasoning_effort": "high"  # Request the high reasoning budget

    # "enable_thinking": False  # Optional: disable all thinking

}

response = requests.post(url, headers=headers, data=json.dumps(payload))
print(response.json()["choices"][0]["message"]["content"])

```

### curl Command

For quick testing or shell scripts, pass the parameter directly in the JSON payload:

```bash
curl https://api.z.ai/v1/chat/completions \
  -H "Authorization: Bearer $ZHIPU_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
        "model": "glm-5.2",
        "messages": [{"role":"user","content":"Plan a week-long itinerary for a trip to Kyoto."}],
        "reasoning_effort":"high"
      }'

```

### Transformers Client

When using the OpenAI-compatible `transformers` library with local GLM-5 weights, pass `reasoning_effort` as a generation argument:

```python
from transformers import AutoModelForCausalLM, AutoTokenizer

tokenizer = AutoTokenizer.from_pretrained("zai-org/glm-5.2")
model = AutoModelForCausalLM.from_pretrained("zai-org/glm-5.2")

inputs = tokenizer(
    [{"role": "user", "content": "Write a short poem about autumn."}],
    return_tensors="pt",
    padding=True,
)

outputs = model.generate(
    **inputs,
    max_new_tokens=256,
    do_sample=True,
    reasoning_effort="high"  # Custom generation argument

)

print(tokenizer.decode(outputs[0], skip_special_tokens=True))

```

## Source Code Reference

The implementation details are documented across several key files in the `zai-org/GLM-5` repository:

- **[`README.md`](https://github.com/zai-org/GLM-5/blob/main/README.md)** – Documents the `reasoning_effort` parameter defaults and behavior descriptions.
- **[`README_zh.md`](https://github.com/zai-org/GLM-5/blob/main/README_zh.md)** – Localized Chinese documentation covering the same reasoning controls.
- **[`example/ascend.md`](https://github.com/zai-org/GLM-5/blob/main/example/ascend.md)** – Deployment guide for Ascend NPU platforms, including optional inference flags for reasoning budget.
- **[`requirements.txt`](https://github.com/zai-org/GLM-5/blob/main/requirements.txt)** – Lists Python dependencies required for client libraries that support these parameters.

## Summary

- **Use `reasoning_effort: "max"`** for default, latency-optimized responses suitable for most production workloads.
- **Use `reasoning_effort: "high"`** when you need deeper Chain-of-Thought reasoning, complex problem solving, or higher accuracy at the cost of increased latency.
- **Set `enable_thinking: false`** to disable reasoning entirely for maximum speed in real-time applications.
- The parameter is validated at the API gateway and passed to the inference engine without requiring changes to other request fields.
- Internal mechanisms include Dynamic Sparse Attention budget allocation and iterative tool call scheduling.

## Frequently Asked Questions

### What is the difference between max and high reasoning effort in GLM-5?

**`max`** provides a bounded thinking budget that stops after a preset number of internal steps, optimizing for speed while maintaining strong performance. **`high`** removes these constraints, allowing the model to perform more Chain-of-Thought iterations and access broader context windows, resulting in more thorough reasoning but slower response times.

### Can I disable reasoning entirely to get faster responses?

Yes. Set **`enable_thinking=false`** in your API request to completely disable the internal CoT reasoning mechanism. This bypasses the thinking token and returns the fastest possible answer, which is useful for latency-critical real-time applications or very low-resource environments.

### Does changing reasoning_effort require modifying the model weights?

No. The `reasoning_effort` parameter is a runtime inference flag parsed by the API gateway and passed to the inference engine. You can switch between `"max"` and `"high"` without reloading model weights or changing your deployment configuration, as the control operates on the inference budget rather than the model architecture itself.

### Which GLM-5 model versions support the reasoning_effort parameter?

The parameter is supported across the GLM-5 family including **`glm-5`**, **`glm-5.1`**, and **`glm-5.2`**. The implementation is consistent across these versions, though newer models may exhibit more sophisticated internal reasoning patterns when the `"high"` setting is activated.