Controlling Reasoning Effort in GLM-5 API: max vs high
The GLM-5 API provides a reasoning_effort parameter that lets you trade latency for deeper reasoning by setting "max" (default, faster) or "high" (more thorough Chain-of-Thought).
The zai-org/GLM-5 repository exposes a request-level control for reasoning depth that directly impacts how the GLM-5, GLM-5.2, and GLM-5.1 models allocate internal compute resources. By adjusting the reasoning_effort parameter, developers can optimize for either low-latency responses or more accurate, deeply-reasoned outputs without modifying model weights or deployment infrastructure.
Understanding the reasoning_effort Parameter
The reasoning_effort parameter accepts string values that determine the budget for internal Chain-of-Thought (CoT) processing. According to the source documentation in README.md, this parameter is parsed at the API gateway and passed downstream to the inference engine.
max (Default)
"max" is the default setting that activates a standard reasoning budget optimized for speed-to-answer. When this value is set, the model uses a bounded thinking token that stops after a preset number of internal steps, guaranteeing consistent latency for general inference tasks. This level is ideal for benchmark reproductions, production workloads, and real-time applications where response time is critical.
high
"high" instructs the model to allocate a larger reasoning budget, keeping the "thinking" token active for significantly more steps. In this mode, the transformer decoder performs more extensive internal Chain-of-Thought steps and searches the knowledge store more thoroughly. This produces more accurate or creative answers but increases latency due to additional compute allocation. Use this for complex coding problems, multi-step planning tasks, or when the default output lacks sufficient depth.
Disabling Thinking Entirely
Set enable_thinking=false to completely disable the model's internal CoT reasoning. This returns the fastest possible answer by bypassing the thinking token entirely, making it suitable for real-time UI feedback or extremely low-resource environments where any reasoning overhead is unacceptable.
How Reasoning Effort Works Internally
When the request arrives at the serving layer, the JSON payload is parsed and the reasoning_effort value triggers specific mechanisms in the inference engine. As implemented in zai-org/GLM-5, the high setting lifts the thinking budget flag, which influences two primary subsystems:
-
Sparse Attention Budget – The Dynamic Sparse Attention (DSA) module allocates additional compute to the "thinking" token, allowing it to attend to more context windows and maintain longer coherent reasoning chains.
-
Iterative Tool Calls – In agentic modes, the model may schedule extra tool calls (such as code execution or search operations) before finalizing its response. The
"high"level permits more such iterations, while the default"max"level enforces stricter limits to maintain performance guarantees.
The parameter validation occurs in the API gateway logic before being passed to the model inference process, ensuring no additional parameters need adjustment when switching between modes.
Implementation Examples
Python with Requests
Use the Z.ai API endpoint to control reasoning effort programmatically:
import requests
import json
url = "https://api.z.ai/v1/chat/completions"
headers = {
"Authorization": f"Bearer {YOUR_ZHIPU_API_KEY}",
"Content-Type": "application/json"
}
payload = {
"model": "glm-5.2", # or glm-5.1 / glm-5
"messages": [
{"role": "user", "content": "Explain the difference between recursion and iteration."}
],
"reasoning_effort": "high" # Request the high reasoning budget
# "enable_thinking": False # Optional: disable all thinking
}
response = requests.post(url, headers=headers, data=json.dumps(payload))
print(response.json()["choices"][0]["message"]["content"])
curl Command
For quick testing or shell scripts, pass the parameter directly in the JSON payload:
curl https://api.z.ai/v1/chat/completions \
-H "Authorization: Bearer $ZHIPU_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "glm-5.2",
"messages": [{"role":"user","content":"Plan a week-long itinerary for a trip to Kyoto."}],
"reasoning_effort":"high"
}'
Transformers Client
When using the OpenAI-compatible transformers library with local GLM-5 weights, pass reasoning_effort as a generation argument:
from transformers import AutoModelForCausalLM, AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("zai-org/glm-5.2")
model = AutoModelForCausalLM.from_pretrained("zai-org/glm-5.2")
inputs = tokenizer(
[{"role": "user", "content": "Write a short poem about autumn."}],
return_tensors="pt",
padding=True,
)
outputs = model.generate(
**inputs,
max_new_tokens=256,
do_sample=True,
reasoning_effort="high" # Custom generation argument
)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
Source Code Reference
The implementation details are documented across several key files in the zai-org/GLM-5 repository:
README.md– Documents thereasoning_effortparameter defaults and behavior descriptions.README_zh.md– Localized Chinese documentation covering the same reasoning controls.example/ascend.md– Deployment guide for Ascend NPU platforms, including optional inference flags for reasoning budget.requirements.txt– Lists Python dependencies required for client libraries that support these parameters.
Summary
- Use
reasoning_effort: "max"for default, latency-optimized responses suitable for most production workloads. - Use
reasoning_effort: "high"when you need deeper Chain-of-Thought reasoning, complex problem solving, or higher accuracy at the cost of increased latency. - Set
enable_thinking: falseto disable reasoning entirely for maximum speed in real-time applications. - The parameter is validated at the API gateway and passed to the inference engine without requiring changes to other request fields.
- Internal mechanisms include Dynamic Sparse Attention budget allocation and iterative tool call scheduling.
Frequently Asked Questions
What is the difference between max and high reasoning effort in GLM-5?
max provides a bounded thinking budget that stops after a preset number of internal steps, optimizing for speed while maintaining strong performance. high removes these constraints, allowing the model to perform more Chain-of-Thought iterations and access broader context windows, resulting in more thorough reasoning but slower response times.
Can I disable reasoning entirely to get faster responses?
Yes. Set enable_thinking=false in your API request to completely disable the internal CoT reasoning mechanism. This bypasses the thinking token and returns the fastest possible answer, which is useful for latency-critical real-time applications or very low-resource environments.
Does changing reasoning_effort require modifying the model weights?
No. The reasoning_effort parameter is a runtime inference flag parsed by the API gateway and passed to the inference engine. You can switch between "max" and "high" without reloading model weights or changing your deployment configuration, as the control operates on the inference budget rather than the model architecture itself.
Which GLM-5 model versions support the reasoning_effort parameter?
The parameter is supported across the GLM-5 family including glm-5, glm-5.1, and glm-5.2. The implementation is consistent across these versions, though newer models may exhibit more sophisticated internal reasoning patterns when the "high" setting is activated.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →