LLM Reasoning Tokens in Speech-to-Speech: How They Work and How to Disable Them
When using OpenAI-compatible backends in the Speech-to-Speech pipeline, reasoning tokens can be suppressed by configuring the extra_body HTTP parameter via either the disable_thinking flag or the reasoning_effort argument depending on your provider.
The Hugging Face Speech-to-Speech repository supports language models that generate internal chain-of-thought reasoning before producing final output. For real-time speech applications, these intermediate reasoning tokens can introduce unwanted latency and content. Understanding how to disable LLM reasoning tokens ensures your pipeline returns only the user-intended speech content.
How Reasoning Tokens Are Generated
When the Speech-to-Speech pipeline initializes an OpenAI-compatible handler, it prepares an HTTP payload that may include provider-specific flags controlling whether the model "thinks" before responding.
In src/speech_to_speech/LLM/base_openai_compatible_language_model.py, the setup() method captures configuration flags during initialization (lines 137-141):
# From src/speech_to_speech/LLM/base_openai_compatible_language_model.py lines 137-141
self._extra_body = {}
if self.args.disable_thinking:
self._extra_body["chat_template_kwargs"] = {"enable_thinking": False}
if self.args.reasoning_effort:
self._extra_body["reasoning_effort"] = self.args.reasoning_effort
The private method _build_extra_body() (lines 185-196) constructs the final payload dictionary. When reasoning_effort is non-empty, it adds {"reasoning_effort": <value>}. Otherwise, when disable_thinking is enabled (the default), it sends {"chat_template_kwargs": {"enable_thinking": False}}.
For official OpenAI API endpoints, this method returns None (lines 192-194), meaning reasoning tokens are only generated when connecting to third-party compatible backends that support thinking modes.
Disabling Reasoning Tokens: Two Methods
The repository provides two mutually exclusive approaches to turn off reasoning tokens, each targeting different provider implementations.
Method 1: The disable_thinking Flag (Default)
disable_thinking=True sends chat_template_kwargs.enable_thinking=False to providers that respect this parameter, including vLLM and Qwen deployments.
This is configured in ResponsesApiLanguageModelHandlerArguments (lines 28-34 of src/speech_to_speech/arguments_classes/responses_api_language_model_arguments.py):
handler = ResponsesApiLanguageModelHandler(
model_name="qwen2.5-7b-instruct",
responses_api_base_url="https://your-vllm-endpoint.com/v1",
responses_api_api_key="YOUR_KEY",
responses_api_disable_thinking=True, # Disables reasoning tokens
)
Method 2: The reasoning_effort Argument
reasoning_effort="none" overrides the disable_thinking flag and sends the reasoning_effort key instead. This is required for providers like GLM via the Hugging Face router that ignore chat-template flags.
Defined in ChatCompletionsLanguageModelArguments and utilized in handlers:
handler = ResponsesApiLanguageModelHandler(
model_name="glm-4-0620",
responses_api_base_url="https://hf.co/api",
responses_api_api_key="YOUR_KEY",
responses_api_reasoning_effort="none", # Disables reasoning for GLM
)
| Method | Parameter | Payload Sent | Supported Providers |
|---|---|---|---|
| Chat Template Override | disable_thinking=True |
{"chat_template_kwargs": {"enable_thinking": False}} |
vLLM, Qwen |
| Effort Level | reasoning_effort="none" |
{"reasoning_effort": "none"} |
GLM, compatible routers |
Provider-Specific Implementation Details
Different inference providers interpret reasoning controls differently. The extra_body payload is only attached to non-official OpenAI endpoints, as the official API handles reasoning through separate mechanisms.
- vLLM and Qwen: Respect the
chat_template_kwargs.enable_thinking=falseflag sent whendisable_thinking=True. - GLM via HF Router: Ignores the chat-template flag and requires
reasoning_effort='none'to suppress reasoning. - Official OpenAI: The
_build_extra_body()method returnsNonefor official endpoints (lines 192-194), so these flags have no effect onapi.openai.com.
The extra body configuration is ultimately forwarded to the HTTP client in responses_api_language_model.py (line 102) and chat_completions_language_model.py (line 104).
Configuration via CLI and Programmatic API
Both disabling methods are exposed through argument classes for flexible configuration.
Using the CLI:
python -m speech_to_speech.main \
--responses_api_base_url https://api.together.xyz/v1 \
--responses_api_api_key $API_KEY \
--responses_api_disable_thinking # Disables reasoning tokens
Programmatic configuration with explicit reasoning effort:
from speech_to_speech.LLM.responses_api_language_model import ResponsesApiLanguageModelHandler
# For providers requiring reasoning_effort
handler = ResponsesApiLanguageModelHandler(
model_name="glm-4",
responses_api_base_url="https://hf.co/api",
responses_api_api_key="key",
responses_api_reasoning_effort="none"
)
Unit tests in tests/test_responses_api_language_model.py and tests/test_chat_completions_backend.py verify that these configurations correctly populate the extra body payload.
Summary
- Reasoning tokens are chain-of-thought outputs generated by certain LLMs before the final response, controlled via the
extra_bodyHTTP field in OpenAI-compatible backends. disable_thinking=True(default) sendschat_template_kwargs.enable_thinking=Falseand works with vLLM and Qwen providers.reasoning_effort="none"sends a specific effort level and is required for providers like GLM that ignore chat-template flags.- Official OpenAI endpoints receive
Nonefor extra body, making these flags ineffective forapi.openai.com. - Configuration is available both programmatically through handler arguments and via the CLI using
--responses_api_disable_thinking.
Frequently Asked Questions
What are reasoning tokens in the context of Speech-to-Speech?
Reasoning tokens are intermediate chain-of-thought outputs generated by large language models (such as Qwen or GLM) before producing the final response. In the Speech-to-Speech pipeline, these tokens appear in the HTTP response stream but are typically hidden from the user; however, they add latency and computational overhead. The pipeline routes these through OpenAI-compatible handlers as implemented in base_openai_compatible_language_model.py.
Why does the official OpenAI API ignore the reasoning token settings?
According to lines 192-194 in src/speech_to_speech/LLM/base_openai_compatible_language_model.py, the _build_extra_body() method returns None when the endpoint is the official OpenAI API (api.openai.com). This design prevents sending unsupported parameters to OpenAI's native endpoints while allowing third-party compatible providers (vLLM, Together, GLM) to receive provider-specific reasoning controls through the extra_body field.
How do I know which disable method my provider supports?
If you are using vLLM or Qwen deployments, use disable_thinking=True (the default). If you are using GLM via the Hugging Face router or similar providers that do not respect chat-template flags, you must set reasoning_effort="none" instead. Check your provider's OpenAI-compatible API documentation to determine whether they use the chat_template_kwargs or reasoning_effort parameter for controlling chain-of-thought generation.
Can I enable reasoning tokens if I want the model to show its work?
Yes. Set disable_thinking=False when constructing your handler, and ensure reasoning_effort is empty or set to a value like "low", "medium", or "high". This removes the suppression flags from the extra_body payload, allowing the backend to generate and return reasoning tokens if the underlying model supports chain-of-thought modes.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →