How to Switch Between Max and High Thinking Effort Modes in GLM-5.2
To switch between max and high thinking effort modes in GLM-5.2, set the reasoning_effort parameter to "high" for reduced latency, or omit it (or use any other value) for the default max mode which provides exhaustive chain-of-thought reasoning.
GLM-5.2, the coding-focused variant of the GLM-5 series from the zai-org/GLM-5 repository, exposes a reasoning_effort parameter that controls how much computational budget the model allocates to chain-of-thought generation. This parameter determines whether the model runs in max effort mode for thorough reasoning or high effort mode for faster responses.
Understanding the Reasoning Effort Parameter
According to the repository documentation in README.md at line 81, the reasoning_effort parameter accepts two distinct values:
max(default): The model performs exhaustive chain-of-thought generation using the full reasoning budget. This mode activates automatically when the parameter is omitted or set to any value other than"high".high: The model uses a reduced reasoning budget, trading a small amount of performance for lower latency.
Implementation Examples
Using the ZhipuAI Client
The OpenAI-compatible zhipuai client supports the reasoning_effort parameter directly in the chat completions API:
from zhipuai import ZhipuAI
client = ZhipuAI(api_key="YOUR_API_KEY")
# Max effort (default) - exhaustive reasoning
resp_max = client.chat.completions.create(
model="glm-5.2",
messages=[{"role": "user", "content": "Write a quick sort implementation in Python"}]
)
print(resp_max.choices[0].message.content)
# High effort - reduced latency
resp_high = client.chat.completions.create(
model="glm-5.2",
messages=[{"role": "user", "content": "Write a quick sort implementation in Python"}],
reasoning_effort="high" # Explicit switch to high effort
)
print(resp_high.choices[0].message.content)
Configuring vLLM for Local Inference
When serving GLM-5.2 locally with vLLM, pass the reasoning_effort parameter through the SamplingParams class:
from vllm import LLM, SamplingParams
llm = LLM(model="zai-org/GLM-5.2")
# Max effort (default behavior)
sampler_max = SamplingParams(temperature=0.7)
# High effort mode
sampler_high = SamplingParams(
temperature=0.7,
reasoning_effort="high"
)
# Generate with max effort
output_max = llm.generate(
prompt="Explain the decorator pattern.",
sampling_params=sampler_max
)
print(output_max[0].outputs[0].text)
# Generate with high effort
output_high = llm.generate(
prompt="Explain the decorator pattern.",
sampling_params=sampler_high
)
print(output_high[0].outputs[0].text)
SGLang YAML Configuration
For SGLang deployments, specify the effort level in your YAML recipe:
# glm5_high_effort.yaml
model: "glm-5.2"
prompt: |
<|User|>
Generate a unit test suite for the following function:
def add(a, b): return a + b
<|Assistant|>
reasoning_effort: "high"
Execute the configuration with:
sglang run glm5_high_effort.yaml
Disabling Chain-of-Thought Completely
To bypass reasoning entirely and receive direct answers without chain-of-thought generation, use the enable_thinking parameter set to false:
resp = client.chat.completions.create(
model="glm-5.2",
messages=[{"role": "user", "content": "Summarize the plot of Moby-Dick"}],
enable_thinking=False # Disables thinking, overrides reasoning_effort
)
This parameter is documented alongside reasoning_effort in the repository's README files.
Summary
- The
reasoning_effortparameter in GLM-5.2 controls the computational budget allocated to reasoning, withmaxas the default for exhaustive chain-of-thought andhighfor reduced latency. - To switch to high effort mode, explicitly pass
reasoning_effort="high"in your API request; any other value or omission defaults to max effort. - The parameter works across multiple serving backends including the ZhipuAI client, vLLM, and SGLang.
- To disable reasoning entirely, set
enable_thinking=False. - Configuration details are documented in
README.mdandREADME_zh.mdat line 81 in the zai-org/GLM-5 repository.
Frequently Asked Questions
What is the default thinking effort mode in GLM-5.2?
The default mode is max effort, which provides exhaustive chain-of-thought generation. This mode activates automatically when the reasoning_effort parameter is omitted or set to any value other than "high", as implemented in the inference API documented in the repository's README.
When should I use high effort mode instead of max?
Use high effort mode when you need reduced latency and can tolerate slightly less thorough reasoning. This setting reduces the computational budget for chain-of-thought generation, making it ideal for rapid prototyping or high-throughput scenarios where response speed matters more than exhaustive analysis.
Can I disable thinking entirely in GLM-5.2?
Yes. Set the enable_thinking parameter to false to receive direct answers without any chain-of-thought generation. This overrides the reasoning_effort setting and is useful when you need straightforward responses without intermediate reasoning steps.
Does the reasoning_effort parameter work with all serving backends?
The parameter is supported across the major serving frameworks mentioned in the zai-org/GLM-5 repository, including the official ZhipuAI client, vLLM, and SGLang. Always verify your specific backend version supports the parameter, as documented in example/ascend.md for specialized deployments like Ascend NPU.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →