How to Use the reasoning_effort Parameter to Control Thinking Levels in GLM-5.2
Set reasoning_effort="high" to enable deep, thorough reasoning on complex queries, omit the field entirely for default max mode (faster, lower latency), or pass enable_thinking=false to disable internal chain-of-thought completely.
The GLM-5.2 model (also referred to as GLM-S.2) in the zai-org/GLM-5 repository exposes a reasoning_effort parameter that acts as a thinking budget controller. This API field allows you to trade latency for reasoning depth, making it straightforward to optimize the model's behavior for everything from quick completions to multi-step mathematical proofs. According to the repository's documentation in README.md (lines 81–82), the parameter supports three distinct operational modes.
Understanding the reasoning_effort Values
The reasoning_effort field accepts string values that map directly to internal compute budgets. When you send a request to the chat completions endpoint, the model checks this field to determine how much chain-of-thought processing to apply before generating a response.
Default Mode (max)
When you omit the reasoning_effort field from your payload, or set it to "max", the model operates in its default fast-reasoning mode. This configuration provides normal-level reasoning with reduced latency, making it ideal for benchmark reproduction, simple Q&A, or high-throughput applications. The max value represents the standard thinking budget and is automatically applied when the parameter is missing.
High-Effort Mode (high)
Passing reasoning_effort="high" instructs the model to allocate significantly more compute to internal reasoning steps. This mode produces deeper, more thorough analysis but introduces higher latency and increased token consumption. Use this setting for complex mathematical proofs, intricate logic puzzles, or tasks requiring multi-step deduction.
Disabling Chain-of-Thought (enable_thinking)
To completely suppress the model's internal reasoning stream, add "enable_thinking": false to your request payload. This flag overrides the reasoning_effort budget and treats the model as a standard completion engine, removing all chain-of-thought-style processing. This is useful for pure token-generation tasks where intermediate reasoning steps are unnecessary.
Code Examples
Below are production-ready implementations using Python with the requests library and standard cURL commands. All examples target the https://api.z.ai/v1/chat/completions endpoint using the glm-s-2 model identifier.
Default (Fast) Mode
Omitting reasoning_effort yields the fastest response times:
import requests
import json
payload = {
"model": "glm-s-2",
"messages": [{"role": "user", "content": "Explain the difference between recursion and iteration."}]
# No reasoning_effort field → defaults to max
}
response = requests.post(
"https://api.z.ai/v1/chat/completions",
headers={
"Authorization": f"Bearer {YOUR_API_KEY}",
"Content-Type": "application/json"
},
data=json.dumps(payload)
)
print(response.json())
High-Effort Reasoning
Explicitly set the parameter for complex problem-solving:
payload = {
"model": "glm-s-2",
"messages": [{"role": "user", "content": "Explain quantum entanglement in simple terms."}],
"reasoning_effort": "high"
}
response = requests.post(
"https://api.z.ai/v1/chat/completions",
headers={
"Authorization": f"Bearer {YOUR_API_KEY}",
"Content-Type": "application/json"
},
data=json.dumps(payload)
)
Disable Thinking Completely
Suppress all internal reasoning with the boolean flag:
payload = {
"model": "glm-s-2",
"messages": [{"role": "user", "content": "Generate a list of 5 random numbers."}],
"reasoning_effort": "max",
"enable_thinking": False
}
response = requests.post(
"https://api.z.ai/v1/chat/completions",
headers={
"Authorization": f"Bearer {YOUR_API_KEY}",
"Content-Type": "application/json"
},
data=json.dumps(payload)
)
cURL Equivalents
For shell-based integrations:
# Default (max) mode
curl -X POST https://api.z.ai/v1/chat/completions \
-H "Authorization: Bearer $ZHIPU_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "glm-s-2",
"messages": [{"role":"user","content":"What is the capital of France?"}]
}'
# High-effort reasoning
curl -X POST https://api.z.ai/v1/chat/completions \
-H "Authorization: Bearer $ZHIPU_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "glm-s-2",
"messages": [{"role":"user","content":"Explain quantum entanglement in simple terms."}],
"reasoning_effort": "high"
}'
# Disable thinking
curl -X POST https://api.z.ai/v1/chat/completions \
-H "Authorization: Bearer $ZHIPU_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "glm-s-2",
"messages": [{"role":"user","content":"Generate a list of 5 random numbers."}],
"reasoning_effort": "max",
"enable_thinking": false
}'
Source Code Reference
The behavior documented above is hardcoded in the repository's configuration documentation:
README.md(lines 81–82): Contains the official description of thereasoning_effortandenable_thinkingflags, explicitly stating thatmaxis the default andhighrequires explicit configuration.README_zh.md(line 81): Provides the Chinese-language equivalent of the same guidance for bilingual implementations.example/ascend.md: Demonstrates hardware-specific deployment contexts where you might combinereasoning_effortsettings with specialized Ascend NPU configurations.
Summary
- Omit
reasoning_effortto run GLM-5.2 in defaultmaxmode for fast, low-latency responses. - Set
reasoning_effort="high"when you need deep, thorough reasoning on complex tasks, accepting the trade-off of higher latency. - Use
enable_thinking=falseto completely disable chain-of-thought processing, treating the model as a standard completion engine. - All GLM-5 family models share this API contract, allowing consistent reasoning control across the model ecosystem.
Frequently Asked Questions
What is the default value for reasoning_effort in GLM-5.2?
The default value is max. When you omit the reasoning_effort field from your request payload entirely, or pass any value other than high, the model automatically defaults to max mode. This provides normal-level reasoning with optimized latency, as documented in README.md lines 81–82.
Does setting reasoning_effort to high increase token usage?
Yes. The high setting instructs the model to perform more thorough internal reasoning before generating a response, which typically consumes additional tokens and increases latency. While the output quality improves for complex tasks, you should expect higher compute costs and slower response times compared to the default max mode.
Can I disable chain-of-thought reasoning while keeping reasoning_effort set to max?
Yes. You can combine "reasoning_effort": "max" with "enable_thinking": false in the same payload. The enable_thinking boolean acts as an override that suppresses all internal reasoning steps, regardless of the budget allocated by reasoning_effort. This configuration is useful when you want fast completions without any intermediate processing steps.
Is the reasoning_effort parameter available in all GLM-5 family models?
Yes. According to the source analysis, GLM-S.2 follows the same API contract as the other GLM-5 family models. The reasoning_effort parameter is a standardized field across the model family, meaning you can use identical request payloads when switching between different GLM-5 variants, including when deploying on specialized hardware as shown in example/ascend.md.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →