# Differences Between responses-api and chat-completions LLM Backends in Speech-to-Speech

> Understand responses-api vs chat-completions LLM backends. Choose the right one for basic streaming or advanced tool-call streaming and reasoning controls.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: deep-dive
- Published: 2026-07-30

---

**The responses-api backend communicates via the legacy `/v1/responses` endpoint with basic streaming support, while the chat-completions backend uses the modern `/v1/chat/completions` protocol featuring robust tool-call streaming and extended reasoning controls.**

The Hugging Face speech-to-speech (STS) repository provides two interchangeable OpenAI-compatible HTTP endpoints for language model inference. Understanding the architectural differences between the responses-api and chat-completions LLM backends allows you to optimize your voice-agent pipeline for specific deployment scenarios, from legacy server compatibility to advanced tool-calling with self-hosted vLLM instances.

## Core Architectural Differences

### Endpoint Specifications and Handler Classes

The two backends implement distinct handler classes that target different API endpoints. The **responses-api** backend uses `ResponsesApiModelHandler` located in [`src/speech_to_speech/LLM/responses_api_language_model.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/responses_api_language_model.py), which defaults to sending requests to `POST /v1/responses`. Conversely, the **chat-completions** backend implements `ChatCompletionsApiModelHandler` in [`src/speech_to_speech/LLM/chat_completions_language_model.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/chat_completions_language_model.py), targeting the `POST /v1/chat/completions` endpoint.

Both handlers share a common argument foundation through `ResponsesApiLanguageModelHandlerArguments` defined in [`src/speech_to_speech/arguments_classes/responses_api_language_model_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/responses_api_language_model_arguments.py). This base class provides CLI flags such as `--responses_api_base_url`, `--responses_api_api_key`, and `--responses_api_stream` that work identically across both backends.

### Message Conversion and Payload Adaptation

The chat-completions backend performs critical message transformation through the `_chat_messages` method (lines 55-66 in [`chat_completions_language_model.py`](https://github.com/huggingface/speech-to-speech/blob/main/chat_completions_language_model.py)). This conversion logic performs two essential adaptations:

- **Tool Call Serialization**: Ensures `tool_calls.arguments` are properly formatted as JSON strings rather than raw objects
- **Multimodal Rewriting**: Transforms Realtime API format elements (`input_text`, `input_image`) into Chat-Completions compatible shapes (`text`, `image_url`)

The responses-api backend sends chat payloads as-is without transformation, maintaining the native OpenAI Realtime format.

### Streaming Protocol Implementation

Both backends implement an `_iter_stream_events` method to process streaming responses, but handle different response schemas:

- **responses-api**: Consumes `Stream[ResponseChunk]` objects and extracts tool-call deltas from the Responses-specific format
- **chat-completions**: Processes `Stream[ChatCompletionChunk]` objects using the stricter Chat-Completions schema, which provides more reliable delta parsing for complex tool invocations

Despite these differences, both handlers ultimately emit identical internal events—including `AssistantMessage`, `ToolCall`, `TextDelta`, and `Usage`—ensuring the rest of the STS pipeline (VAD → STT → TTS) remains backend-agnostic.

## Feature Comparison: Tool Calling and Reasoning Control

### Tool-Call Reliability and Format

The **responses-api** backend uses `ResponseFunctionToolCall` objects with streaming deltas arriving in `choices[].delta.tool_calls`. While functional, this format occasionally exhibits compatibility issues with certain server implementations.

The **chat-completions** backend leverages the mature Chat-Completions streaming protocol, which has proven more reliable for vLLM and Qwen tool-calling scenarios (see repository issue #312). This protocol handles nested tool-call arguments more robustly during streaming inference.

### Reasoning and Thinking Parameters

Both backends respect the `responses_api_disable_thinking` flag to control model reasoning behavior. However, **chat-completions** extends this capability through an additional parameter defined in `ChatCompletionsLanguageModelHandlerArguments` at [`src/speech_to_speech/arguments_classes/chat_completions_language_model_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/chat_completions_language_model_arguments.py).

The `responses_api_reasoning_effort` parameter allows you to send `extra_body={'reasoning_effort': <value>}` to force providers to skip or reduce reasoning effort when they ignore the generic `disable_thinking` flag. This proves essential when deploying models through certain vLLM configurations that require explicit reasoning controls in the request body.

### Warm-up and Connection Handling

Each backend implements a `warmup` method to validate connectivity before processing real-time audio:

- **responses-api**: Sends a minimal request to `/v1/responses`
- **chat-completions**: Sends a minimal request to `/v1/chat/completions` (see lines 94-102 in [`chat_completions_language_model.py`](https://github.com/huggingface/speech-to-speech/blob/main/chat_completions_language_model.py))

These warm-up routines ensure the endpoint is responsive before the voice conversation begins, preventing cold-start latency during active sessions.

## When to Use Each Backend

Select **responses-api** when you require broad compatibility with older OpenAI-compatible providers or prefer the simplest streaming format without message conversion overhead. This backend works reliably with any server implementing the legacy Responses specification.

Choose **chat-completions** for the following scenarios:

- **Self-hosted vLLM or llama.cpp servers** where tool-call streaming proves more stable through the Chat-Completions endpoint rather than the Responses API
- **Advanced reasoning control** when you need the `responses_api_reasoning_effort` knob to suppress thinking on providers that bypass the standard `disable_thinking` flag
- **Production tool-calling** with complex multimodal inputs requiring the robust payload adaptation provided by the `_chat_messages` conversion layer

## Configuration and Code Examples

### CLI Configuration

Activate the responses-api backend (default) using standard flags:

```bash
speech-to-speech \
    --mode realtime \
    --stt parakeet-tdt \
    --llm_backend responses-api \
    --tts qwen3 \
    --model_name "gpt-4o-mini" \
    --responses_api_api_key "$OPENAI_API_KEY" \
    --responses_api_stream

```

Switch to the chat-completions backend with extended reasoning controls:

```bash
speech-to-speech \
    --mode realtime \
    --stt parakeet-tdt \
    --llm_backend chat-completions \
    --tts qwen3 \
    --model_name "Qwen/Qwen3-4B-Instruct-2507" \
    --responses_api_base_url "http://localhost:8000/v1" \
    --responses_api_stream \
    --responses_api_reasoning_effort none

```

### Python Client Integration

Direct OpenAI client usage with the responses-api endpoint:

```python
from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:8000/v1",
    api_key="any-string",
)

response = client.responses.create(
    model="gpt-4o-mini",
    messages=[{"role": "user", "content": "Hello"}],
    stream=True,
)

```

Direct OpenAI client usage with the chat-completions endpoint and reasoning control:

```python
from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:8000/v1",
    api_key="any-string",
)

response = client.chat.completions.create(
    model="gpt-4o-mini",
    messages=[{"role": "user", "content": "Hello"}],
    stream=True,
    extra_body={"reasoning_effort": "none"},
)

```

### Pipeline Selection

The backend selection occurs during pipeline construction in [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py) (lines 104-110), where the `--llm_backend` flag value determines which handler class instantiates the LLM module. Both handlers accept identical base configuration parameters, simplifying backend swaps without modifying downstream audio processing components.

## Summary

- **Endpoint Differences**: responses-api targets `/v1/responses` while chat-completions uses `/v1/chat/completions`
- **Handler Classes**: `ResponsesApiModelHandler` vs `ChatCompletionsApiModelHandler` in their respective files under `src/speech_to_speech/LLM/`
- **Message Processing**: chat-completions converts payloads via `_chat_messages` to ensure JSON string arguments and proper multimodal formatting
- **Tool Calling**: chat-completions provides more reliable streaming for vLLM and Qwen tool invocations
- **Reasoning Control**: chat-completions adds `responses_api_reasoning_effort` via `ChatCompletionsLanguageModelHandlerArguments` for providers ignoring standard disable flags
- **Pipeline Compatibility**: Both backends emit identical internal events, ensuring seamless integration with STT and TTS components

## Frequently Asked Questions

### What is the default LLM backend in speech-to-speech?

The **responses-api** backend serves as the default when you omit the `--llm_backend` flag. This default is enforced during argument parsing in the pipeline construction phase, ensuring backward compatibility with existing deployments that expect the legacy `/v1/responses` endpoint behavior.

### Can I switch between backends without modifying my pipeline code?

Yes. Both backends emit identical internal event types (`AssistantMessage`, `ToolCall`, `Usage`), making the rest of the pipeline agnostic to your choice. You only need to change the `--llm_backend` flag and potentially adjust the `--responses_api_reasoning_effort` parameter when switching to chat-completions.

### Why does chat-completions have an extra reasoning_effort parameter?

The `responses_api_reasoning_effort` parameter exists in `ChatCompletionsLanguageModelHandlerArguments` to address providers (such as certain vLLM versions) that ignore the generic `disable_thinking` flag. It sends reasoning controls via `extra_body`, forcing the model to skip or reduce reasoning effort when standard flags fail.

### Which backend should I use with self-hosted vLLM servers?

Use **chat-completions** when running self-hosted vLLM or llama.cpp instances, particularly if you rely on tool-calling functionality. The Chat-Completions streaming protocol handles tool-call deltas more reliably than the Responses API format, resolving compatibility issues documented in repository issue #312.