# How to Enable or Disable Thinking Mode in GLM-5: Complete Configuration Guide

> Learn how to enable or disable thinking mode in GLM-5 using enable_thinking and reasoning_effort. Configure GLM-5 for optimal output quality and inference speed.

- Repository: [Z.ai/GLM-5](https://github.com/zai-org/GLM-5)
- Tags: how-to-guide
- Published: 2026-06-19

---

**GLM-5 provides two parameters to control reasoning behavior: `reasoning_effort` (set to `"max"` or `"high"`) and `enable_thinking` (boolean `true` or `false`), allowing you to balance between output quality and inference latency.**

GLM-5, developed by the zai-org/GLM-5 repository, introduces configurable thinking modes that let developers control the depth of internal reasoning. By adjusting the `reasoning_effort` and `enable_thinking` parameters, you can optimize the model for everything from complex reasoning tasks to latency-critical applications. This guide covers the exact implementation details found in the source documentation and provides practical code examples for each configuration.

## Understanding GLM-5's Thinking Mode Controls

GLM-5 implements two distinct parameters that govern its reasoning pipeline. According to the documentation in [`/README.md`](https://github.com/zai-org/GLM-5/blob/main//README.md) (line 80) and [`/README_zh.md`](https://github.com/zai-org/GLM-5/blob/main//README_zh.md) (line 80), these controls allow precise tuning of the model's cognitive depth.

### The reasoning_effort Parameter

The `reasoning_effort` parameter selects the amount of internal reasoning the model performs:

- **`"max"`** (default): Maximum reasoning effort, recommended for most workloads including benchmarking and standard inference
- **`"high"`**: More compute-intensive reasoning level for complex tasks requiring deeper analysis

When omitted or set to any value other than `"high"`, the model defaults to **Max** mode.

### The enable_thinking Parameter

The `enable_thinking` parameter acts as a master switch for the reasoning stage:

- **`true`** (default): Enables the reasoning pipeline (implicit when parameter is omitted)
- **`false`**: Completely disables thinking, overriding any `reasoning_effort` setting

This boolean control is particularly useful for latency-sensitive applications where immediate responses are required.

## Configuration Reference

The following table summarizes the parameter specifications as implemented in the GLM-5 source code:

| Parameter | Purpose | Accepted Values | Default |
|-----------|---------|----------------|---------|
| `reasoning_effort` | Selects reasoning depth | `"max"`, `"high"` | `"max"` |
| `enable_thinking` | Toggles reasoning on/off | `true`, `false` | `true` |

## Implementation Examples

Below are concrete implementations using the GLM-5 Python client. These patterns demonstrate how to configure thinking mode for different use cases based on the repository's example code.

### Default Max Reasoning

To use the default maximum reasoning mode, simply omit both parameters or explicitly accept the defaults:

```python
response = glm_client.generate(
    prompt="Explain the difference between supervised and unsupervised learning.",
    # No reasoning_effort or enable_thinking supplied → defaults to Max

)
print(response.text)

```

### High-Effort Reasoning

For complex architectural or mathematical tasks requiring deeper analysis:

```python
response = glm_client.generate(
    prompt="Design a scalable microservice architecture for an e-commerce platform.",
    reasoning_effort="high"  # Explicitly request high-effort reasoning

)
print(response.text)

```

### Disabling Thinking for Low Latency

When response speed is critical and reasoning is unnecessary:

```python
response = glm_client.generate(
    prompt="What is 2 + 2?",
    enable_thinking=False  # Skip reasoning stage entirely

)
print(response.text)  # Returns fast, direct answer

```

### Combining Parameters

Note that `enable_thinking` takes precedence over `reasoning_effort`. When disabled, reasoning is skipped regardless of the effort level:

```python

# enable_thinking=False overrides reasoning_effort

response = glm_client.generate(
    prompt="Summarize the latest advances in quantum computing.",
    reasoning_effort="high",
    enable_thinking=False
)
print(response.text)

```

## Source Code Documentation

The thinking mode parameters are documented in the following files within the zai-org/GLM-5 repository:

- **[`/README.md`](https://github.com/zai-org/GLM-5/blob/main//README.md)** (line 80): Primary English documentation describing parameter behavior
- **[`/README_zh.md`](https://github.com/zai-org/GLM-5/blob/main//README_zh.md)** (line 80): Chinese version with equivalent parameter details
- **[`example/ascend.md`](https://github.com/zai-org/GLM-5/blob/main/example/ascend.md)**: Usage scenarios and implementation examples
- **[`skills/glm-master-s/SKILL.md`](https://github.com/zai-org/GLM-5/blob/main/skills/glm-master-s/SKILL.md)**: Skill definition file outlining platform integration

## Summary

- **GLM-5** exposes two controls: `reasoning_effort` for depth selection and `enable_thinking` for toggling reasoning entirely
- Default configuration uses `"max"` reasoning effort with thinking enabled implicitly
- Set `reasoning_effort="high"` for computationally intensive reasoning tasks
- Set `enable_thinking=False` to disable reasoning completely and reduce latency
- The boolean `enable_thinking` parameter overrides any `reasoning_effort` setting when set to `false`

## Frequently Asked Questions

### What is the default thinking mode in GLM-5?

By default, GLM-5 operates with `reasoning_effort` set to `"max"` and `enable_thinking` implicitly set to `true`. This configuration provides maximum reasoning capability without requiring explicit parameter configuration, as documented in [`/README.md`](https://github.com/zai-org/GLM-5/blob/main//README.md).

### Can I use high reasoning effort with thinking disabled?

While you can technically pass both parameters simultaneously, setting `enable_thinking=False` disables the reasoning stage entirely, overriding any `reasoning_effort` value passed. The model will skip internal reasoning regardless of the effort level requested.

### Which parameter should I use to reduce API latency?

To minimize latency, set `enable_thinking=False`. This bypasses the internal reasoning stage completely, causing the model to generate responses immediately without the computational overhead of the reasoning pipeline.

### Where are these parameters documented in the source code?

The thinking mode parameters are documented at line 80 in both [`/README.md`](https://github.com/zai-org/GLM-5/blob/main//README.md) and [`/README_zh.md`](https://github.com/zai-org/GLM-5/blob/main//README_zh.md), with additional implementation examples available in [`example/ascend.md`](https://github.com/zai-org/GLM-5/blob/main/example/ascend.md) and the skill definition in [`skills/glm-master-s/SKILL.md`](https://github.com/zai-org/GLM-5/blob/main/skills/glm-master-s/SKILL.md).