# LLM Reasoning Tokens in Speech-to-Speech: How They Work and How to Disable Them

> Understand how LLM reasoning tokens impact speech-to-speech and learn to disable them using the disable thinking flag or reasoning effort argument in your OpenAI-compatible backend.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: internals
- Published: 2026-07-30

---

**When using OpenAI-compatible backends in the Speech-to-Speech pipeline, reasoning tokens can be suppressed by configuring the `extra_body` HTTP parameter via either the `disable_thinking` flag or the `reasoning_effort` argument depending on your provider.**

The Hugging Face Speech-to-Speech repository supports language models that generate internal chain-of-thought reasoning before producing final output. For real-time speech applications, these intermediate reasoning tokens can introduce unwanted latency and content. Understanding how to disable LLM reasoning tokens ensures your pipeline returns only the user-intended speech content.

## How Reasoning Tokens Are Generated

When the Speech-to-Speech pipeline initializes an OpenAI-compatible handler, it prepares an HTTP payload that may include provider-specific flags controlling whether the model "thinks" before responding.

In [`src/speech_to_speech/LLM/base_openai_compatible_language_model.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/base_openai_compatible_language_model.py), the `setup()` method captures configuration flags during initialization (lines 137-141):

```python

# From src/speech_to_speech/LLM/base_openai_compatible_language_model.py lines 137-141

self._extra_body = {}
if self.args.disable_thinking:
    self._extra_body["chat_template_kwargs"] = {"enable_thinking": False}
if self.args.reasoning_effort:
    self._extra_body["reasoning_effort"] = self.args.reasoning_effort

```

The private method `_build_extra_body()` (lines 185-196) constructs the final payload dictionary. When `reasoning_effort` is non-empty, it adds `{"reasoning_effort": <value>}`. Otherwise, when `disable_thinking` is enabled (the default), it sends `{"chat_template_kwargs": {"enable_thinking": False}}`.

For official OpenAI API endpoints, this method returns `None` (lines 192-194), meaning reasoning tokens are only generated when connecting to third-party compatible backends that support thinking modes.

## Disabling Reasoning Tokens: Two Methods

The repository provides two mutually exclusive approaches to turn off reasoning tokens, each targeting different provider implementations.

### Method 1: The `disable_thinking` Flag (Default)

**`disable_thinking=True`** sends `chat_template_kwargs.enable_thinking=False` to providers that respect this parameter, including **vLLM** and **Qwen** deployments.

This is configured in `ResponsesApiLanguageModelHandlerArguments` (lines 28-34 of [`src/speech_to_speech/arguments_classes/responses_api_language_model_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/responses_api_language_model_arguments.py)):

```python
handler = ResponsesApiLanguageModelHandler(
    model_name="qwen2.5-7b-instruct",
    responses_api_base_url="https://your-vllm-endpoint.com/v1",
    responses_api_api_key="YOUR_KEY",
    responses_api_disable_thinking=True,  # Disables reasoning tokens

)

```

### Method 2: The `reasoning_effort` Argument

**`reasoning_effort="none"`** overrides the `disable_thinking` flag and sends the `reasoning_effort` key instead. This is required for providers like **GLM via the Hugging Face router** that ignore chat-template flags.

Defined in `ChatCompletionsLanguageModelArguments` and utilized in handlers:

```python
handler = ResponsesApiLanguageModelHandler(
    model_name="glm-4-0620",
    responses_api_base_url="https://hf.co/api",
    responses_api_api_key="YOUR_KEY",
    responses_api_reasoning_effort="none",  # Disables reasoning for GLM

)

```

| Method | Parameter | Payload Sent | Supported Providers |
|--------|-----------|--------------|---------------------|
| **Chat Template Override** | `disable_thinking=True` | `{"chat_template_kwargs": {"enable_thinking": False}}` | vLLM, Qwen |
| **Effort Level** | `reasoning_effort="none"` | `{"reasoning_effort": "none"}` | GLM, compatible routers |

## Provider-Specific Implementation Details

Different inference providers interpret reasoning controls differently. The `extra_body` payload is only attached to non-official OpenAI endpoints, as the official API handles reasoning through separate mechanisms.

- **vLLM and Qwen**: Respect the `chat_template_kwargs.enable_thinking=false` flag sent when `disable_thinking=True`.
- **GLM via HF Router**: Ignores the chat-template flag and requires `reasoning_effort='none'` to suppress reasoning.
- **Official OpenAI**: The `_build_extra_body()` method returns `None` for official endpoints (lines 192-194), so these flags have no effect on `api.openai.com`.

The extra body configuration is ultimately forwarded to the HTTP client in [`responses_api_language_model.py`](https://github.com/huggingface/speech-to-speech/blob/main/responses_api_language_model.py) (line 102) and [`chat_completions_language_model.py`](https://github.com/huggingface/speech-to-speech/blob/main/chat_completions_language_model.py) (line 104).

## Configuration via CLI and Programmatic API

Both disabling methods are exposed through argument classes for flexible configuration.

**Using the CLI:**

```bash
python -m speech_to_speech.main \
    --responses_api_base_url https://api.together.xyz/v1 \
    --responses_api_api_key $API_KEY \
    --responses_api_disable_thinking  # Disables reasoning tokens

```

**Programmatic configuration with explicit reasoning effort:**

```python
from speech_to_speech.LLM.responses_api_language_model import ResponsesApiLanguageModelHandler

# For providers requiring reasoning_effort

handler = ResponsesApiLanguageModelHandler(
    model_name="glm-4",
    responses_api_base_url="https://hf.co/api",
    responses_api_api_key="key",
    responses_api_reasoning_effort="none"
)

```

Unit tests in [`tests/test_responses_api_language_model.py`](https://github.com/huggingface/speech-to-speech/blob/main/tests/test_responses_api_language_model.py) and [`tests/test_chat_completions_backend.py`](https://github.com/huggingface/speech-to-speech/blob/main/tests/test_chat_completions_backend.py) verify that these configurations correctly populate the extra body payload.

## Summary

- **Reasoning tokens** are chain-of-thought outputs generated by certain LLMs before the final response, controlled via the `extra_body` HTTP field in OpenAI-compatible backends.
- **`disable_thinking=True`** (default) sends `chat_template_kwargs.enable_thinking=False` and works with vLLM and Qwen providers.
- **`reasoning_effort="none"`** sends a specific effort level and is required for providers like GLM that ignore chat-template flags.
- **Official OpenAI endpoints** receive `None` for extra body, making these flags ineffective for `api.openai.com`.
- Configuration is available both programmatically through handler arguments and via the CLI using `--responses_api_disable_thinking`.

## Frequently Asked Questions

### What are reasoning tokens in the context of Speech-to-Speech?

Reasoning tokens are intermediate chain-of-thought outputs generated by large language models (such as Qwen or GLM) before producing the final response. In the Speech-to-Speech pipeline, these tokens appear in the HTTP response stream but are typically hidden from the user; however, they add latency and computational overhead. The pipeline routes these through OpenAI-compatible handlers as implemented in [`base_openai_compatible_language_model.py`](https://github.com/huggingface/speech-to-speech/blob/main/base_openai_compatible_language_model.py).

### Why does the official OpenAI API ignore the reasoning token settings?

According to lines 192-194 in [`src/speech_to_speech/LLM/base_openai_compatible_language_model.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/base_openai_compatible_language_model.py), the `_build_extra_body()` method returns `None` when the endpoint is the official OpenAI API (`api.openai.com`). This design prevents sending unsupported parameters to OpenAI's native endpoints while allowing third-party compatible providers (vLLM, Together, GLM) to receive provider-specific reasoning controls through the `extra_body` field.

### How do I know which disable method my provider supports?

If you are using **vLLM** or **Qwen** deployments, use `disable_thinking=True` (the default). If you are using **GLM via the Hugging Face router** or similar providers that do not respect chat-template flags, you must set `reasoning_effort="none"` instead. Check your provider's OpenAI-compatible API documentation to determine whether they use the `chat_template_kwargs` or `reasoning_effort` parameter for controlling chain-of-thought generation.

### Can I enable reasoning tokens if I want the model to show its work?

Yes. Set `disable_thinking=False` when constructing your handler, and ensure `reasoning_effort` is empty or set to a value like `"low"`, `"medium"`, or `"high"`. This removes the suppression flags from the `extra_body` payload, allowing the backend to generate and return reasoning tokens if the underlying model supports chain-of-thought modes.