How to Enable or Disable Thinking Mode in GLM-5: Complete Configuration Guide

GLM-5 provides two parameters to control reasoning behavior: reasoning_effort (set to "max" or "high") and enable_thinking (boolean true or false), allowing you to balance between output quality and inference latency.

GLM-5, developed by the zai-org/GLM-5 repository, introduces configurable thinking modes that let developers control the depth of internal reasoning. By adjusting the reasoning_effort and enable_thinking parameters, you can optimize the model for everything from complex reasoning tasks to latency-critical applications. This guide covers the exact implementation details found in the source documentation and provides practical code examples for each configuration.

Understanding GLM-5's Thinking Mode Controls

GLM-5 implements two distinct parameters that govern its reasoning pipeline. According to the documentation in /README.md (line 80) and /README_zh.md (line 80), these controls allow precise tuning of the model's cognitive depth.

The reasoning_effort Parameter

The reasoning_effort parameter selects the amount of internal reasoning the model performs:

  • "max" (default): Maximum reasoning effort, recommended for most workloads including benchmarking and standard inference
  • "high": More compute-intensive reasoning level for complex tasks requiring deeper analysis

When omitted or set to any value other than "high", the model defaults to Max mode.

The enable_thinking Parameter

The enable_thinking parameter acts as a master switch for the reasoning stage:

  • true (default): Enables the reasoning pipeline (implicit when parameter is omitted)
  • false: Completely disables thinking, overriding any reasoning_effort setting

This boolean control is particularly useful for latency-sensitive applications where immediate responses are required.

Configuration Reference

The following table summarizes the parameter specifications as implemented in the GLM-5 source code:

Parameter Purpose Accepted Values Default
reasoning_effort Selects reasoning depth "max", "high" "max"
enable_thinking Toggles reasoning on/off true, false true

Implementation Examples

Below are concrete implementations using the GLM-5 Python client. These patterns demonstrate how to configure thinking mode for different use cases based on the repository's example code.

Default Max Reasoning

To use the default maximum reasoning mode, simply omit both parameters or explicitly accept the defaults:

response = glm_client.generate(
    prompt="Explain the difference between supervised and unsupervised learning.",
    # No reasoning_effort or enable_thinking supplied → defaults to Max

)
print(response.text)

High-Effort Reasoning

For complex architectural or mathematical tasks requiring deeper analysis:

response = glm_client.generate(
    prompt="Design a scalable microservice architecture for an e-commerce platform.",
    reasoning_effort="high"  # Explicitly request high-effort reasoning

)
print(response.text)

Disabling Thinking for Low Latency

When response speed is critical and reasoning is unnecessary:

response = glm_client.generate(
    prompt="What is 2 + 2?",
    enable_thinking=False  # Skip reasoning stage entirely

)
print(response.text)  # Returns fast, direct answer

Combining Parameters

Note that enable_thinking takes precedence over reasoning_effort. When disabled, reasoning is skipped regardless of the effort level:


# enable_thinking=False overrides reasoning_effort

response = glm_client.generate(
    prompt="Summarize the latest advances in quantum computing.",
    reasoning_effort="high",
    enable_thinking=False
)
print(response.text)

Source Code Documentation

The thinking mode parameters are documented in the following files within the zai-org/GLM-5 repository:

Summary

  • GLM-5 exposes two controls: reasoning_effort for depth selection and enable_thinking for toggling reasoning entirely
  • Default configuration uses "max" reasoning effort with thinking enabled implicitly
  • Set reasoning_effort="high" for computationally intensive reasoning tasks
  • Set enable_thinking=False to disable reasoning completely and reduce latency
  • The boolean enable_thinking parameter overrides any reasoning_effort setting when set to false

Frequently Asked Questions

What is the default thinking mode in GLM-5?

By default, GLM-5 operates with reasoning_effort set to "max" and enable_thinking implicitly set to true. This configuration provides maximum reasoning capability without requiring explicit parameter configuration, as documented in /README.md.

Can I use high reasoning effort with thinking disabled?

While you can technically pass both parameters simultaneously, setting enable_thinking=False disables the reasoning stage entirely, overriding any reasoning_effort value passed. The model will skip internal reasoning regardless of the effort level requested.

Which parameter should I use to reduce API latency?

To minimize latency, set enable_thinking=False. This bypasses the internal reasoning stage completely, causing the model to generate responses immediately without the computational overhead of the reasoning pipeline.

Where are these parameters documented in the source code?

The thinking mode parameters are documented at line 80 in both /README.md and /README_zh.md, with additional implementation examples available in example/ascend.md and the skill definition in skills/glm-master-s/SKILL.md.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →