How to Enable or Disable Thinking Mode in GLM-5: Complete Configuration Guide
GLM-5 provides two parameters to control reasoning behavior: reasoning_effort (set to "max" or "high") and enable_thinking (boolean true or false), allowing you to balance between output quality and inference latency.
GLM-5, developed by the zai-org/GLM-5 repository, introduces configurable thinking modes that let developers control the depth of internal reasoning. By adjusting the reasoning_effort and enable_thinking parameters, you can optimize the model for everything from complex reasoning tasks to latency-critical applications. This guide covers the exact implementation details found in the source documentation and provides practical code examples for each configuration.
Understanding GLM-5's Thinking Mode Controls
GLM-5 implements two distinct parameters that govern its reasoning pipeline. According to the documentation in /README.md (line 80) and /README_zh.md (line 80), these controls allow precise tuning of the model's cognitive depth.
The reasoning_effort Parameter
The reasoning_effort parameter selects the amount of internal reasoning the model performs:
"max"(default): Maximum reasoning effort, recommended for most workloads including benchmarking and standard inference"high": More compute-intensive reasoning level for complex tasks requiring deeper analysis
When omitted or set to any value other than "high", the model defaults to Max mode.
The enable_thinking Parameter
The enable_thinking parameter acts as a master switch for the reasoning stage:
true(default): Enables the reasoning pipeline (implicit when parameter is omitted)false: Completely disables thinking, overriding anyreasoning_effortsetting
This boolean control is particularly useful for latency-sensitive applications where immediate responses are required.
Configuration Reference
The following table summarizes the parameter specifications as implemented in the GLM-5 source code:
| Parameter | Purpose | Accepted Values | Default |
|---|---|---|---|
reasoning_effort |
Selects reasoning depth | "max", "high" |
"max" |
enable_thinking |
Toggles reasoning on/off | true, false |
true |
Implementation Examples
Below are concrete implementations using the GLM-5 Python client. These patterns demonstrate how to configure thinking mode for different use cases based on the repository's example code.
Default Max Reasoning
To use the default maximum reasoning mode, simply omit both parameters or explicitly accept the defaults:
response = glm_client.generate(
prompt="Explain the difference between supervised and unsupervised learning.",
# No reasoning_effort or enable_thinking supplied → defaults to Max
)
print(response.text)
High-Effort Reasoning
For complex architectural or mathematical tasks requiring deeper analysis:
response = glm_client.generate(
prompt="Design a scalable microservice architecture for an e-commerce platform.",
reasoning_effort="high" # Explicitly request high-effort reasoning
)
print(response.text)
Disabling Thinking for Low Latency
When response speed is critical and reasoning is unnecessary:
response = glm_client.generate(
prompt="What is 2 + 2?",
enable_thinking=False # Skip reasoning stage entirely
)
print(response.text) # Returns fast, direct answer
Combining Parameters
Note that enable_thinking takes precedence over reasoning_effort. When disabled, reasoning is skipped regardless of the effort level:
# enable_thinking=False overrides reasoning_effort
response = glm_client.generate(
prompt="Summarize the latest advances in quantum computing.",
reasoning_effort="high",
enable_thinking=False
)
print(response.text)
Source Code Documentation
The thinking mode parameters are documented in the following files within the zai-org/GLM-5 repository:
/README.md(line 80): Primary English documentation describing parameter behavior/README_zh.md(line 80): Chinese version with equivalent parameter detailsexample/ascend.md: Usage scenarios and implementation examplesskills/glm-master-s/SKILL.md: Skill definition file outlining platform integration
Summary
- GLM-5 exposes two controls:
reasoning_effortfor depth selection andenable_thinkingfor toggling reasoning entirely - Default configuration uses
"max"reasoning effort with thinking enabled implicitly - Set
reasoning_effort="high"for computationally intensive reasoning tasks - Set
enable_thinking=Falseto disable reasoning completely and reduce latency - The boolean
enable_thinkingparameter overrides anyreasoning_effortsetting when set tofalse
Frequently Asked Questions
What is the default thinking mode in GLM-5?
By default, GLM-5 operates with reasoning_effort set to "max" and enable_thinking implicitly set to true. This configuration provides maximum reasoning capability without requiring explicit parameter configuration, as documented in /README.md.
Can I use high reasoning effort with thinking disabled?
While you can technically pass both parameters simultaneously, setting enable_thinking=False disables the reasoning stage entirely, overriding any reasoning_effort value passed. The model will skip internal reasoning regardless of the effort level requested.
Which parameter should I use to reduce API latency?
To minimize latency, set enable_thinking=False. This bypasses the internal reasoning stage completely, causing the model to generate responses immediately without the computational overhead of the reasoning pipeline.
Where are these parameters documented in the source code?
The thinking mode parameters are documented at line 80 in both /README.md and /README_zh.md, with additional implementation examples available in example/ascend.md and the skill definition in skills/glm-master-s/SKILL.md.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →