# Difference Between responses-api and chat-completions LLM Backends in Speech-to-Speech

> Understand the difference between responses-api and chat-completions LLM backends in speech-to-speech on Hugging Face. Learn about legacy vs modern API features and tool call streaming.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: deep-dive
- Published: 2026-07-09

---

**The `responses-api` backend connects to OpenAI's legacy `/v1/responses` endpoint, while `chat-completions` uses the modern `/v1/chat/completions` API with enhanced tool-call streaming and message format conversion.**

The Hugging Face Speech-to-Speech (STS) pipeline supports two interchangeable OpenAI-compatible HTTP backends for Large Language Model (LLM) inference. Understanding the difference between `responses-api` and `chat-completions` LLM backends helps you select the right protocol for your deployment, whether you need legacy compatibility with older providers or robust tool-calling support with self-hosted servers.

## What Are the LLM Backends?

Speech-to-Speech can communicate with language models through two distinct HTTP endpoints that share a common argument foundation. Both backends inherit from **`ResponsesApiLanguageModelHandlerArguments`** defined in [`src/speech_to_speech/arguments_classes/responses_api_language_model_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/responses_api_language_model_arguments.py), which provides CLI flags like `--responses_api_base_url`, `--responses_api_api_key`, and `--responses_api_stream`.

The **`chat-completions`** backend extends these arguments through **`ChatCompletionsLanguageModelHandlerArguments`** in [`src/speech_to_speech/arguments_classes/chat_completions_language_model_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/chat_completions_language_model_arguments.py), adding the unique `responses_api_reasoning_effort` parameter for providers that require explicit reasoning control.

## Key Architectural Differences

### Endpoint and Handler Classes

Each backend uses a distinct endpoint and handler implementation:

- **`responses-api`**: Sends requests to `POST /v1/responses` via the `ResponsesApiModelHandler` class in [`src/speech_to_speech/LLM/responses_api_language_model.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/responses_api_language_model.py).
- **`chat-completions`**: Sends requests to `POST /v1/chat/completions` via the `ChatCompletionsApiModelHandler` class in [`src/speech_to_speech/LLM/chat_completions_language_model.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/chat_completions_language_model.py).

### Message Format Conversion

The **Chat-Completions** backend performs explicit message adaptation before sending requests. In [`src/speech_to_speech/LLM/chat_completions_language_model.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/chat_completions_language_model.py) (lines 55-66), the `_chat_messages` method:

- Ensures `tool_calls.arguments` are serialized as JSON strings rather than raw objects.
- Rewrites multimodal content from Realtime API shapes (`input_text`, `input_image`) to Chat-Completions format (`text`, `image_url`).

The **Responses-API** backend sends the chat payload as-is in OpenAI Realtime format without conversion.

### Tool-Call Streaming

Both handlers implement `_iter_stream_events` to process streaming responses, but they handle different response schemas:

- **Responses-API**: Gathers tool-call deltas from `Stream[ResponseChunk]` objects using the older `ResponseFunctionToolCall` format.
- **Chat-Completions**: Processes `Stream[ChatCompletionChunk]` objects (lines 5-12 in [`chat_completions_language_model.py`](https://github.com/huggingface/speech-to-speech/blob/main/chat_completions_language_model.py)), emitting `AssistantMessage`, `TextDelta`, and `ToolCall` events using the stricter Chat-Completions schema that is more reliable for vLLM and Qwen tool-calling.

### Reasoning Control

The backends differ in how they handle reasoning or "thinking" modes:

- **Responses-API**: Controlled solely by the `responses_api_disable_thinking` flag.
- **Chat-Completions**: Supports both `responses_api_disable_thinking` and the additional `responses_api_reasoning_effort` parameter, which sends `extra_body={'reasoning_effort': <value>}` to force providers to skip or reduce reasoning effort when the generic flag is ignored.

## When to Use Each Backend

Use **`responses-api`** when you need compatibility with older OpenAI-compatible providers or simple deployments where the legacy streaming format is sufficient.

Use **`chat-completions`** when:

- Running self-hosted vLLM or llama.cpp servers where the Chat-Completions endpoint provides more reliable tool-call streaming (see issue #312).
- You need the `responses_api_reasoning_effort` knob to suppress reasoning on providers that ignore the standard `disable_thinking` flag.
- Working with multimodal inputs that require conversion from Realtime API formats to standard Chat-Completions shapes.

## Configuration and CLI Examples

Select your backend using the `--llm_backend` flag parsed in [`module_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/module_arguments.py) and injected into the pipeline in [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py) (lines 104-110).

**Using the Responses-API backend (default):**

```bash
speech-to-speech \
    --mode realtime \
    --stt parakeet-tdt \
    --llm_backend responses-api \
    --tts qwen3 \
    --model_name "gpt-4o-mini" \
    --responses_api_api_key "$OPENAI_API_KEY" \
    --responses_api_stream

```

**Switching to the Chat-Completions backend:**

```bash
speech-to-speech \
    --mode realtime \
    --stt parakeet-tdt \
    --llm_backend chat-completions \
    --tts qwen3 \
    --model_name "Qwen/Qwen3-4B-Instruct-2507" \
    --responses_api_base_url "http://localhost:8000/v1" \
    --responses_api_stream \
    --responses_api_reasoning_effort none

```

**Direct OpenAI client usage (Responses-API):**

```python
from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:8000/v1",
    api_key="any-string",
)

response = client.responses.create(
    model="gpt-4o-mini",
    messages=[{"role": "user", "content": "Hello"}],
    stream=True,
)

```

**Direct OpenAI client usage (Chat-Completions):**

```python
from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:8000/v1",
    api_key="any-string",
)

response = client.chat.completions.create(
    model="gpt-4o-mini",
    messages=[{"role": "user", "content": "Hello"}],
    stream=True,
    extra_body={"reasoning_effort": "none"},
)

```

## Summary

- Both backends inherit from `ResponsesApiLanguageModelHandlerArguments` and share base CLI flags, but `chat-completions` adds the `responses_api_reasoning_effort` parameter.
- **Responses-API** uses `/v1/responses` with `ResponsesApiModelHandler` and sends messages in Realtime format without conversion.
- **Chat-Completions** uses `/v1/chat/completions` with `ChatCompletionsApiModelHandler`, converting tool arguments to JSON strings and rewriting multimodal content.
- The Chat-Completions backend provides more reliable tool-call streaming for vLLM and Qwen deployments.
- Select backends via `--llm_backend` (default: `responses-api`) in [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py).

## Frequently Asked Questions

### Which backend should I use with vLLM?

Use the **`chat-completions`** backend when running vLLM servers. According to the source code in [`chat_completions_language_model.py`](https://github.com/huggingface/speech-to-speech/blob/main/chat_completions_language_model.py), this backend provides more reliable tool-call streaming compared to the Responses-API format, which can be flaky with certain vLLM configurations (see issue #312).

### How do I disable reasoning or "thinking" in the LLM?

For the **Responses-API** backend, use the `--responses_api_disable_thinking` flag. For the **Chat-Completions** backend, you can use both `--responses_api_disable_thinking` and `--responses_api_reasoning_effort none`. The latter sends `extra_body={'reasoning_effort': 'none'}` to providers that ignore the standard disable flag.

### Are the backends interchangeable in the pipeline?

Yes. Both handlers ultimately produce the same internal events (`AssistantMessage`, `ToolCall`, `Usage`, etc.), making the rest of the pipeline (VAD → STT → TTS) agnostic to which backend you choose. The selection only affects the HTTP protocol and message formatting layer.

### What is the default LLM backend?

The default is **`responses-api`**. You can override this by setting `--llm_backend chat-completions` in your CLI arguments, which the pipeline constructor reads from [`module_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/module_arguments.py) and applies in [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py) during handler initialization.