Difference Between responses-api and chat-completions LLM Backends in Speech-to-Speech
The responses-api backend connects to OpenAI's legacy /v1/responses endpoint, while chat-completions uses the modern /v1/chat/completions API with enhanced tool-call streaming and message format conversion.
The Hugging Face Speech-to-Speech (STS) pipeline supports two interchangeable OpenAI-compatible HTTP backends for Large Language Model (LLM) inference. Understanding the difference between responses-api and chat-completions LLM backends helps you select the right protocol for your deployment, whether you need legacy compatibility with older providers or robust tool-calling support with self-hosted servers.
What Are the LLM Backends?
Speech-to-Speech can communicate with language models through two distinct HTTP endpoints that share a common argument foundation. Both backends inherit from ResponsesApiLanguageModelHandlerArguments defined in src/speech_to_speech/arguments_classes/responses_api_language_model_arguments.py, which provides CLI flags like --responses_api_base_url, --responses_api_api_key, and --responses_api_stream.
The chat-completions backend extends these arguments through ChatCompletionsLanguageModelHandlerArguments in src/speech_to_speech/arguments_classes/chat_completions_language_model_arguments.py, adding the unique responses_api_reasoning_effort parameter for providers that require explicit reasoning control.
Key Architectural Differences
Endpoint and Handler Classes
Each backend uses a distinct endpoint and handler implementation:
responses-api: Sends requests toPOST /v1/responsesvia theResponsesApiModelHandlerclass insrc/speech_to_speech/LLM/responses_api_language_model.py.chat-completions: Sends requests toPOST /v1/chat/completionsvia theChatCompletionsApiModelHandlerclass insrc/speech_to_speech/LLM/chat_completions_language_model.py.
Message Format Conversion
The Chat-Completions backend performs explicit message adaptation before sending requests. In src/speech_to_speech/LLM/chat_completions_language_model.py (lines 55-66), the _chat_messages method:
- Ensures
tool_calls.argumentsare serialized as JSON strings rather than raw objects. - Rewrites multimodal content from Realtime API shapes (
input_text,input_image) to Chat-Completions format (text,image_url).
The Responses-API backend sends the chat payload as-is in OpenAI Realtime format without conversion.
Tool-Call Streaming
Both handlers implement _iter_stream_events to process streaming responses, but they handle different response schemas:
- Responses-API: Gathers tool-call deltas from
Stream[ResponseChunk]objects using the olderResponseFunctionToolCallformat. - Chat-Completions: Processes
Stream[ChatCompletionChunk]objects (lines 5-12 inchat_completions_language_model.py), emittingAssistantMessage,TextDelta, andToolCallevents using the stricter Chat-Completions schema that is more reliable for vLLM and Qwen tool-calling.
Reasoning Control
The backends differ in how they handle reasoning or "thinking" modes:
- Responses-API: Controlled solely by the
responses_api_disable_thinkingflag. - Chat-Completions: Supports both
responses_api_disable_thinkingand the additionalresponses_api_reasoning_effortparameter, which sendsextra_body={'reasoning_effort': <value>}to force providers to skip or reduce reasoning effort when the generic flag is ignored.
When to Use Each Backend
Use responses-api when you need compatibility with older OpenAI-compatible providers or simple deployments where the legacy streaming format is sufficient.
Use chat-completions when:
- Running self-hosted vLLM or llama.cpp servers where the Chat-Completions endpoint provides more reliable tool-call streaming (see issue #312).
- You need the
responses_api_reasoning_effortknob to suppress reasoning on providers that ignore the standarddisable_thinkingflag. - Working with multimodal inputs that require conversion from Realtime API formats to standard Chat-Completions shapes.
Configuration and CLI Examples
Select your backend using the --llm_backend flag parsed in module_arguments.py and injected into the pipeline in src/speech_to_speech/s2s_pipeline.py (lines 104-110).
Using the Responses-API backend (default):
speech-to-speech \
--mode realtime \
--stt parakeet-tdt \
--llm_backend responses-api \
--tts qwen3 \
--model_name "gpt-4o-mini" \
--responses_api_api_key "$OPENAI_API_KEY" \
--responses_api_stream
Switching to the Chat-Completions backend:
speech-to-speech \
--mode realtime \
--stt parakeet-tdt \
--llm_backend chat-completions \
--tts qwen3 \
--model_name "Qwen/Qwen3-4B-Instruct-2507" \
--responses_api_base_url "http://localhost:8000/v1" \
--responses_api_stream \
--responses_api_reasoning_effort none
Direct OpenAI client usage (Responses-API):
from openai import OpenAI
client = OpenAI(
base_url="http://localhost:8000/v1",
api_key="any-string",
)
response = client.responses.create(
model="gpt-4o-mini",
messages=[{"role": "user", "content": "Hello"}],
stream=True,
)
Direct OpenAI client usage (Chat-Completions):
from openai import OpenAI
client = OpenAI(
base_url="http://localhost:8000/v1",
api_key="any-string",
)
response = client.chat.completions.create(
model="gpt-4o-mini",
messages=[{"role": "user", "content": "Hello"}],
stream=True,
extra_body={"reasoning_effort": "none"},
)
Summary
- Both backends inherit from
ResponsesApiLanguageModelHandlerArgumentsand share base CLI flags, butchat-completionsadds theresponses_api_reasoning_effortparameter. - Responses-API uses
/v1/responseswithResponsesApiModelHandlerand sends messages in Realtime format without conversion. - Chat-Completions uses
/v1/chat/completionswithChatCompletionsApiModelHandler, converting tool arguments to JSON strings and rewriting multimodal content. - The Chat-Completions backend provides more reliable tool-call streaming for vLLM and Qwen deployments.
- Select backends via
--llm_backend(default:responses-api) ins2s_pipeline.py.
Frequently Asked Questions
Which backend should I use with vLLM?
Use the chat-completions backend when running vLLM servers. According to the source code in chat_completions_language_model.py, this backend provides more reliable tool-call streaming compared to the Responses-API format, which can be flaky with certain vLLM configurations (see issue #312).
How do I disable reasoning or "thinking" in the LLM?
For the Responses-API backend, use the --responses_api_disable_thinking flag. For the Chat-Completions backend, you can use both --responses_api_disable_thinking and --responses_api_reasoning_effort none. The latter sends extra_body={'reasoning_effort': 'none'} to providers that ignore the standard disable flag.
Are the backends interchangeable in the pipeline?
Yes. Both handlers ultimately produce the same internal events (AssistantMessage, ToolCall, Usage, etc.), making the rest of the pipeline (VAD → STT → TTS) agnostic to which backend you choose. The selection only affects the HTTP protocol and message formatting layer.
What is the default LLM backend?
The default is responses-api. You can override this by setting --llm_backend chat-completions in your CLI arguments, which the pipeline constructor reads from module_arguments.py and applies in s2s_pipeline.py during handler initialization.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →