# Speech-to-Speech responses-api vs chat-completions LLM backends: Key Differences and When to Use Each

> Explore responses-api vs chat-completions LLM backends. Understand key differences in message conversion, reasoning, and streaming to choose the right Speech-to-Speech solution.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: deep-dive
- Published: 2026-07-10

---

**Speech-to-Speech supports two OpenAI-compatible HTTP endpoints—`responses-api` (POST /v1/responses) for legacy streaming compatibility and `chat-completions` (POST /v1/chat/completions) for modern tool-call reliability—with both backends sharing CLI arguments but differing in message conversion, reasoning controls, and streaming protocols.**

The Hugging Face Speech-to-Speech (STS) repository provides dual LLM backend integration, allowing the voice agent pipeline to communicate with language models through either the legacy Responses API or the modern Chat-Completions API. Understanding the architectural distinctions between these **responses-api** and **chat-completions** LLM backends ensures you select the appropriate handler for your specific deployment, whether running self-hosted vLLM instances or connecting to managed OpenAI-compatible providers.

## Core Architectural Differences

### Endpoint URLs and Handler Classes

Each backend maps to a distinct HTTP endpoint and handler implementation within the codebase. The **responses-api** backend utilizes `ResponsesApiModelHandler` defined in [`src/speech_to_speech/LLM/responses_api_language_model.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/responses_api_language_model.py), targeting the endpoint `/v1/responses`. Conversely, the **chat-completions** backend employs `ChatCompletionsApiModelHandler` located in [`src/speech_to_speech/LLM/chat_completions_language_model.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/chat_completions_language_model.py), communicating with `/v1/chat/completions`.

Both handlers ultimately emit identical internal events—`AssistantMessage`, `ToolCall`, `TextDelta`, and `Usage`—ensuring the downstream VAD → STT → TTS pipeline remains agnostic to which transport layer is active. However, the underlying wire protocols and message preprocessing differ significantly.

### Message Format Conversion

The **responses-api** handler transmits conversation history using the raw OpenAI Realtime format, sending messages as-is without transformation. In contrast, the **chat-completions** handler in [`src/speech_to_speech/LLM/chat_completions_language_model.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/chat_completions_language_model.py) implements a `_chat_messages` conversion method (lines 55–66) that performs critical adaptations:

- **Tool call serialization**: Ensures `tool_calls.arguments` are properly formatted as JSON strings rather than raw objects
- **Multimodal translation**: Rewrites Realtime-specific `input_text` and `input_image` fields into Chat-Completions-compatible `text` and `image_url` shapes

This conversion layer makes the chat-completions backend more suitable for providers expecting strict OpenAI Chat-Completions schema compliance.

### Streaming Protocol and Tool-Call Handling

While both backends implement an `_iter_stream_events` method for processing SSE streams, they operate on different chunk types. The **responses-api** handler consumes `Stream[ResponseChunk]` objects and gathers tool-call deltas from `choices[].delta.tool_calls` using the `ResponseFunctionToolCall` format. The **chat-completions** handler processes `Stream[ChatCompletionChunk]` objects with a stricter schema that is more reliable for vLLM + Qwen tool-calling scenarios, addressing known compatibility issues such as vLLM issue #312.

### Reasoning Control Capabilities

Argument inheritance differs subtly between the two. Both backends share the base class `ResponsesApiLanguageModelHandlerArguments` defined in [`src/speech_to_speech/arguments_classes/responses_api_language_model_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/responses_api_language_model_arguments.py), which provides flags like `--responses_api_disable_thinking`. However, the **chat-completions** backend extends this with `ChatCompletionsLanguageModelHandlerArguments` in [`src/speech_to_speech/arguments_classes/chat_completions_language_model_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/chat_completions_language_model_arguments.py), adding the `--responses_api_reasoning_effort` parameter. This knob sends `extra_body={'reasoning_effort': <value>}` to force providers to skip or reduce reasoning effort when the generic `disable_thinking` flag is ignored.

### Warm-up Behavior

Each handler implements distinct warm-up logic to validate connectivity before the real-time conversation begins. The responses-api backend sends minimal validation requests to `/v1/responses`, while the chat-completions implementation targets `/v1/chat/completions` via its `warmup` method (lines 94–102 in [`chat_completions_language_model.py`](https://github.com/huggingface/speech-to-speech/blob/main/chat_completions_language_model.py)).

## When to Use Each Backend

### Use responses-api when:

- You require compatibility with older vLLM builds or providers that only implement the legacy Responses streaming specification
- Your deployment scenario benefits from the simplest, unconverted Realtime format without message transformation overhead
- You do not require the additional `reasoning_effort` control knob for managing model reasoning stages

### Use chat-completions when:

- You are running self-hosted vLLM, llama.cpp, or similar servers where tool-call streaming is more reliable via the Chat-Completions endpoint (particularly relevant for vLLM issue #312)
- You need explicit control over reasoning effort via `--responses_api_reasoning_effort` to suppress thinking on providers that ignore standard disable flags
- Your application requires strict adherence to modern OpenAI Chat-Completions multimodal formatting standards

## CLI and Python Configuration Examples

### Switching Backends via CLI

The backend selection is controlled via the `--llm_backend` flag, parsed in [`module_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/module_arguments.py) and injected during pipeline construction in [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py) (lines 104–110). The default value is `responses-api`.

**Using the default responses-api backend:**

```bash
speech-to-speech \
    --mode realtime \
    --stt parakeet-tdt \
    --llm_backend responses-api \
    --tts qwen3 \
    --model_name "gpt-4o-mini" \
    --responses_api_api_key "$OPENAI_API_KEY" \
    --responses_api_stream

```

**Switching to chat-completions with reasoning control:**

```bash
speech-to-speech \
    --mode realtime \
    --stt parakeet-tdt \
    --llm_backend chat-completions \
    --tts qwen3 \
    --model_name "Qwen/Qwen3-4B-Instruct-2507" \
    --responses_api_base_url "http://localhost:8000/v1" \
    --responses_api_stream \
    --responses_api_reasoning_effort none

```

### Direct Python Client Access

**OpenAI client (responses-api):**

```python
from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:8000/v1",
    api_key="any-string",
)

response = client.responses.create(
    model="gpt-4o-mini",
    messages=[{"role": "user", "content": "Hello"}],
    stream=True,
)

```

**OpenAI client (chat-completions with reasoning override):**

```python
from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:8000/v1",
    api_key="any-string",
)

response = client.chat.completions.create(
    model="gpt-4o-mini",
    messages=[{"role": "user", "content": "Hello"}],
    stream=True,
    extra_body={"reasoning_effort": "none"},
)

```

## Summary

- **responses-api** (`/v1/responses`) uses `ResponsesApiModelHandler` with raw Realtime format streaming, suitable for legacy compatibility and simple deployments
- **chat-completions** (`/v1/chat/completions`) uses `ChatCompletionsApiModelHandler` with message conversion logic (`_chat_messages`) and stricter schema compliance, preferred for modern vLLM tool-calling
- Both backends share base arguments from `ResponsesApiLanguageModelHandlerArguments`, but chat-completions uniquely adds `--responses_api_reasoning_effort` for fine-grained reasoning control
- Selection occurs via `--llm_backend` flag, with pipeline instantiation handled in [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py) at lines 104–110
- Internal event emission remains identical across both backends, ensuring seamless interoperability with the rest of the Speech-to-Speech pipeline

## Frequently Asked Questions

### What CLI flag selects between responses-api and chat-completions?

Use `--llm_backend` followed by either `responses-api` (default) or `chat-completions`. This flag is parsed in [`module_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/module_arguments.py) and determines which handler class—`ResponsesApiModelHandler` or `ChatCompletionsApiModelHandler`—is instantiated during pipeline construction in [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py).

### Do both backends support tool calling in Speech-to-Speech?

Yes. Both handlers generate identical `ToolCall` internal events that the pipeline processes downstream. However, the **chat-completions** backend provides more reliable tool-call streaming for certain self-hosted providers like vLLM, while **responses-api** uses an older delta format that may encounter compatibility issues with specific model implementations.

### Why does chat-completions require message conversion while responses-api does not?

The **chat-completions** backend must transform Realtime API message shapes into strict Chat-Completions schema compliance. Specifically, the `_chat_messages` method in [`chat_completions_language_model.py`](https://github.com/huggingface/speech-to-speech/blob/main/chat_completions_language_model.py) converts multimodal inputs from `input_text`/`input_image` to `text`/`image_url` formats and ensures tool arguments are JSON strings, whereas **responses-api** transmits the native Realtime format directly without transformation.

### Can I use the same API key and base URL for both backends?

Yes. Both backends utilize identical CLI flags for connection parameters (`--responses_api_base_url`, `--responses_api_api_key`) inherited from `ResponsesApiLanguageModelHandlerArguments`. Simply change `--llm_backend` to switch protocols while keeping authentication configuration constant, though ensure the actual endpoint (`/v1/responses` vs `/v1/chat/completions`) is available at the specified base URL.