Speech-to-Speech responses-api vs chat-completions LLM backends: Key Differences and When to Use Each
Speech-to-Speech supports two OpenAI-compatible HTTP endpoints—responses-api (POST /v1/responses) for legacy streaming compatibility and chat-completions (POST /v1/chat/completions) for modern tool-call reliability—with both backends sharing CLI arguments but differing in message conversion, reasoning controls, and streaming protocols.
The Hugging Face Speech-to-Speech (STS) repository provides dual LLM backend integration, allowing the voice agent pipeline to communicate with language models through either the legacy Responses API or the modern Chat-Completions API. Understanding the architectural distinctions between these responses-api and chat-completions LLM backends ensures you select the appropriate handler for your specific deployment, whether running self-hosted vLLM instances or connecting to managed OpenAI-compatible providers.
Core Architectural Differences
Endpoint URLs and Handler Classes
Each backend maps to a distinct HTTP endpoint and handler implementation within the codebase. The responses-api backend utilizes ResponsesApiModelHandler defined in src/speech_to_speech/LLM/responses_api_language_model.py, targeting the endpoint /v1/responses. Conversely, the chat-completions backend employs ChatCompletionsApiModelHandler located in src/speech_to_speech/LLM/chat_completions_language_model.py, communicating with /v1/chat/completions.
Both handlers ultimately emit identical internal events—AssistantMessage, ToolCall, TextDelta, and Usage—ensuring the downstream VAD → STT → TTS pipeline remains agnostic to which transport layer is active. However, the underlying wire protocols and message preprocessing differ significantly.
Message Format Conversion
The responses-api handler transmits conversation history using the raw OpenAI Realtime format, sending messages as-is without transformation. In contrast, the chat-completions handler in src/speech_to_speech/LLM/chat_completions_language_model.py implements a _chat_messages conversion method (lines 55–66) that performs critical adaptations:
- Tool call serialization: Ensures
tool_calls.argumentsare properly formatted as JSON strings rather than raw objects - Multimodal translation: Rewrites Realtime-specific
input_textandinput_imagefields into Chat-Completions-compatibletextandimage_urlshapes
This conversion layer makes the chat-completions backend more suitable for providers expecting strict OpenAI Chat-Completions schema compliance.
Streaming Protocol and Tool-Call Handling
While both backends implement an _iter_stream_events method for processing SSE streams, they operate on different chunk types. The responses-api handler consumes Stream[ResponseChunk] objects and gathers tool-call deltas from choices[].delta.tool_calls using the ResponseFunctionToolCall format. The chat-completions handler processes Stream[ChatCompletionChunk] objects with a stricter schema that is more reliable for vLLM + Qwen tool-calling scenarios, addressing known compatibility issues such as vLLM issue #312.
Reasoning Control Capabilities
Argument inheritance differs subtly between the two. Both backends share the base class ResponsesApiLanguageModelHandlerArguments defined in src/speech_to_speech/arguments_classes/responses_api_language_model_arguments.py, which provides flags like --responses_api_disable_thinking. However, the chat-completions backend extends this with ChatCompletionsLanguageModelHandlerArguments in src/speech_to_speech/arguments_classes/chat_completions_language_model_arguments.py, adding the --responses_api_reasoning_effort parameter. This knob sends extra_body={'reasoning_effort': <value>} to force providers to skip or reduce reasoning effort when the generic disable_thinking flag is ignored.
Warm-up Behavior
Each handler implements distinct warm-up logic to validate connectivity before the real-time conversation begins. The responses-api backend sends minimal validation requests to /v1/responses, while the chat-completions implementation targets /v1/chat/completions via its warmup method (lines 94–102 in chat_completions_language_model.py).
When to Use Each Backend
Use responses-api when:
- You require compatibility with older vLLM builds or providers that only implement the legacy Responses streaming specification
- Your deployment scenario benefits from the simplest, unconverted Realtime format without message transformation overhead
- You do not require the additional
reasoning_effortcontrol knob for managing model reasoning stages
Use chat-completions when:
- You are running self-hosted vLLM, llama.cpp, or similar servers where tool-call streaming is more reliable via the Chat-Completions endpoint (particularly relevant for vLLM issue #312)
- You need explicit control over reasoning effort via
--responses_api_reasoning_effortto suppress thinking on providers that ignore standard disable flags - Your application requires strict adherence to modern OpenAI Chat-Completions multimodal formatting standards
CLI and Python Configuration Examples
Switching Backends via CLI
The backend selection is controlled via the --llm_backend flag, parsed in module_arguments.py and injected during pipeline construction in src/speech_to_speech/s2s_pipeline.py (lines 104–110). The default value is responses-api.
Using the default responses-api backend:
speech-to-speech \
--mode realtime \
--stt parakeet-tdt \
--llm_backend responses-api \
--tts qwen3 \
--model_name "gpt-4o-mini" \
--responses_api_api_key "$OPENAI_API_KEY" \
--responses_api_stream
Switching to chat-completions with reasoning control:
speech-to-speech \
--mode realtime \
--stt parakeet-tdt \
--llm_backend chat-completions \
--tts qwen3 \
--model_name "Qwen/Qwen3-4B-Instruct-2507" \
--responses_api_base_url "http://localhost:8000/v1" \
--responses_api_stream \
--responses_api_reasoning_effort none
Direct Python Client Access
OpenAI client (responses-api):
from openai import OpenAI
client = OpenAI(
base_url="http://localhost:8000/v1",
api_key="any-string",
)
response = client.responses.create(
model="gpt-4o-mini",
messages=[{"role": "user", "content": "Hello"}],
stream=True,
)
OpenAI client (chat-completions with reasoning override):
from openai import OpenAI
client = OpenAI(
base_url="http://localhost:8000/v1",
api_key="any-string",
)
response = client.chat.completions.create(
model="gpt-4o-mini",
messages=[{"role": "user", "content": "Hello"}],
stream=True,
extra_body={"reasoning_effort": "none"},
)
Summary
- responses-api (
/v1/responses) usesResponsesApiModelHandlerwith raw Realtime format streaming, suitable for legacy compatibility and simple deployments - chat-completions (
/v1/chat/completions) usesChatCompletionsApiModelHandlerwith message conversion logic (_chat_messages) and stricter schema compliance, preferred for modern vLLM tool-calling - Both backends share base arguments from
ResponsesApiLanguageModelHandlerArguments, but chat-completions uniquely adds--responses_api_reasoning_effortfor fine-grained reasoning control - Selection occurs via
--llm_backendflag, with pipeline instantiation handled ins2s_pipeline.pyat lines 104–110 - Internal event emission remains identical across both backends, ensuring seamless interoperability with the rest of the Speech-to-Speech pipeline
Frequently Asked Questions
What CLI flag selects between responses-api and chat-completions?
Use --llm_backend followed by either responses-api (default) or chat-completions. This flag is parsed in module_arguments.py and determines which handler class—ResponsesApiModelHandler or ChatCompletionsApiModelHandler—is instantiated during pipeline construction in s2s_pipeline.py.
Do both backends support tool calling in Speech-to-Speech?
Yes. Both handlers generate identical ToolCall internal events that the pipeline processes downstream. However, the chat-completions backend provides more reliable tool-call streaming for certain self-hosted providers like vLLM, while responses-api uses an older delta format that may encounter compatibility issues with specific model implementations.
Why does chat-completions require message conversion while responses-api does not?
The chat-completions backend must transform Realtime API message shapes into strict Chat-Completions schema compliance. Specifically, the _chat_messages method in chat_completions_language_model.py converts multimodal inputs from input_text/input_image to text/image_url formats and ensures tool arguments are JSON strings, whereas responses-api transmits the native Realtime format directly without transformation.
Can I use the same API key and base URL for both backends?
Yes. Both backends utilize identical CLI flags for connection parameters (--responses_api_base_url, --responses_api_api_key) inherited from ResponsesApiLanguageModelHandlerArguments. Simply change --llm_backend to switch protocols while keeping authentication configuration constant, though ensure the actual endpoint (/v1/responses vs /v1/chat/completions) is available at the specified base URL.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →