How to Switch Between Max and High Thinking Effort Modes in GLM-5.2

To switch between max and high thinking effort modes in GLM-5.2, set the reasoning_effort parameter to "high" for reduced latency, or omit it (or use any other value) for the default max mode which provides exhaustive chain-of-thought reasoning.

GLM-5.2, the coding-focused variant of the GLM-5 series from the zai-org/GLM-5 repository, exposes a reasoning_effort parameter that controls how much computational budget the model allocates to chain-of-thought generation. This parameter determines whether the model runs in max effort mode for thorough reasoning or high effort mode for faster responses.

Understanding the Reasoning Effort Parameter

According to the repository documentation in README.md at line 81, the reasoning_effort parameter accepts two distinct values:

  • max (default): The model performs exhaustive chain-of-thought generation using the full reasoning budget. This mode activates automatically when the parameter is omitted or set to any value other than "high".
  • high: The model uses a reduced reasoning budget, trading a small amount of performance for lower latency.

Implementation Examples

Using the ZhipuAI Client

The OpenAI-compatible zhipuai client supports the reasoning_effort parameter directly in the chat completions API:

from zhipuai import ZhipuAI

client = ZhipuAI(api_key="YOUR_API_KEY")

# Max effort (default) - exhaustive reasoning

resp_max = client.chat.completions.create(
    model="glm-5.2",
    messages=[{"role": "user", "content": "Write a quick sort implementation in Python"}]
)
print(resp_max.choices[0].message.content)

# High effort - reduced latency

resp_high = client.chat.completions.create(
    model="glm-5.2",
    messages=[{"role": "user", "content": "Write a quick sort implementation in Python"}],
    reasoning_effort="high"  # Explicit switch to high effort

)
print(resp_high.choices[0].message.content)

Configuring vLLM for Local Inference

When serving GLM-5.2 locally with vLLM, pass the reasoning_effort parameter through the SamplingParams class:

from vllm import LLM, SamplingParams

llm = LLM(model="zai-org/GLM-5.2")

# Max effort (default behavior)

sampler_max = SamplingParams(temperature=0.7)

# High effort mode

sampler_high = SamplingParams(
    temperature=0.7,
    reasoning_effort="high"
)

# Generate with max effort

output_max = llm.generate(
    prompt="Explain the decorator pattern.",
    sampling_params=sampler_max
)
print(output_max[0].outputs[0].text)

# Generate with high effort

output_high = llm.generate(
    prompt="Explain the decorator pattern.",
    sampling_params=sampler_high
)
print(output_high[0].outputs[0].text)

SGLang YAML Configuration

For SGLang deployments, specify the effort level in your YAML recipe:


# glm5_high_effort.yaml

model: "glm-5.2"
prompt: |
  <|User|>
  Generate a unit test suite for the following function:
  def add(a, b): return a + b
  <|Assistant|>
reasoning_effort: "high"

Execute the configuration with:

sglang run glm5_high_effort.yaml

Disabling Chain-of-Thought Completely

To bypass reasoning entirely and receive direct answers without chain-of-thought generation, use the enable_thinking parameter set to false:

resp = client.chat.completions.create(
    model="glm-5.2",
    messages=[{"role": "user", "content": "Summarize the plot of Moby-Dick"}],
    enable_thinking=False  # Disables thinking, overrides reasoning_effort

)

This parameter is documented alongside reasoning_effort in the repository's README files.

Summary

  • The reasoning_effort parameter in GLM-5.2 controls the computational budget allocated to reasoning, with max as the default for exhaustive chain-of-thought and high for reduced latency.
  • To switch to high effort mode, explicitly pass reasoning_effort="high" in your API request; any other value or omission defaults to max effort.
  • The parameter works across multiple serving backends including the ZhipuAI client, vLLM, and SGLang.
  • To disable reasoning entirely, set enable_thinking=False.
  • Configuration details are documented in README.md and README_zh.md at line 81 in the zai-org/GLM-5 repository.

Frequently Asked Questions

What is the default thinking effort mode in GLM-5.2?

The default mode is max effort, which provides exhaustive chain-of-thought generation. This mode activates automatically when the reasoning_effort parameter is omitted or set to any value other than "high", as implemented in the inference API documented in the repository's README.

When should I use high effort mode instead of max?

Use high effort mode when you need reduced latency and can tolerate slightly less thorough reasoning. This setting reduces the computational budget for chain-of-thought generation, making it ideal for rapid prototyping or high-throughput scenarios where response speed matters more than exhaustive analysis.

Can I disable thinking entirely in GLM-5.2?

Yes. Set the enable_thinking parameter to false to receive direct answers without any chain-of-thought generation. This overrides the reasoning_effort setting and is useful when you need straightforward responses without intermediate reasoning steps.

Does the reasoning_effort parameter work with all serving backends?

The parameter is supported across the major serving frameworks mentioned in the zai-org/GLM-5 repository, including the official ZhipuAI client, vLLM, and SGLang. Always verify your specific backend version supports the parameter, as documented in example/ascend.md for specialized deployments like Ascend NPU.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →