Speech-to-Speech responses-api vs chat-completions LLM backends: Key Differences and When to Use Each

Speech-to-Speech supports two OpenAI-compatible HTTP endpoints—responses-api (POST /v1/responses) for legacy streaming compatibility and chat-completions (POST /v1/chat/completions) for modern tool-call reliability—with both backends sharing CLI arguments but differing in message conversion, reasoning controls, and streaming protocols.

The Hugging Face Speech-to-Speech (STS) repository provides dual LLM backend integration, allowing the voice agent pipeline to communicate with language models through either the legacy Responses API or the modern Chat-Completions API. Understanding the architectural distinctions between these responses-api and chat-completions LLM backends ensures you select the appropriate handler for your specific deployment, whether running self-hosted vLLM instances or connecting to managed OpenAI-compatible providers.

Core Architectural Differences

Endpoint URLs and Handler Classes

Each backend maps to a distinct HTTP endpoint and handler implementation within the codebase. The responses-api backend utilizes ResponsesApiModelHandler defined in src/speech_to_speech/LLM/responses_api_language_model.py, targeting the endpoint /v1/responses. Conversely, the chat-completions backend employs ChatCompletionsApiModelHandler located in src/speech_to_speech/LLM/chat_completions_language_model.py, communicating with /v1/chat/completions.

Both handlers ultimately emit identical internal events—AssistantMessage, ToolCall, TextDelta, and Usage—ensuring the downstream VAD → STT → TTS pipeline remains agnostic to which transport layer is active. However, the underlying wire protocols and message preprocessing differ significantly.

Message Format Conversion

The responses-api handler transmits conversation history using the raw OpenAI Realtime format, sending messages as-is without transformation. In contrast, the chat-completions handler in src/speech_to_speech/LLM/chat_completions_language_model.py implements a _chat_messages conversion method (lines 55–66) that performs critical adaptations:

  • Tool call serialization: Ensures tool_calls.arguments are properly formatted as JSON strings rather than raw objects
  • Multimodal translation: Rewrites Realtime-specific input_text and input_image fields into Chat-Completions-compatible text and image_url shapes

This conversion layer makes the chat-completions backend more suitable for providers expecting strict OpenAI Chat-Completions schema compliance.

Streaming Protocol and Tool-Call Handling

While both backends implement an _iter_stream_events method for processing SSE streams, they operate on different chunk types. The responses-api handler consumes Stream[ResponseChunk] objects and gathers tool-call deltas from choices[].delta.tool_calls using the ResponseFunctionToolCall format. The chat-completions handler processes Stream[ChatCompletionChunk] objects with a stricter schema that is more reliable for vLLM + Qwen tool-calling scenarios, addressing known compatibility issues such as vLLM issue #312.

Reasoning Control Capabilities

Argument inheritance differs subtly between the two. Both backends share the base class ResponsesApiLanguageModelHandlerArguments defined in src/speech_to_speech/arguments_classes/responses_api_language_model_arguments.py, which provides flags like --responses_api_disable_thinking. However, the chat-completions backend extends this with ChatCompletionsLanguageModelHandlerArguments in src/speech_to_speech/arguments_classes/chat_completions_language_model_arguments.py, adding the --responses_api_reasoning_effort parameter. This knob sends extra_body={'reasoning_effort': <value>} to force providers to skip or reduce reasoning effort when the generic disable_thinking flag is ignored.

Warm-up Behavior

Each handler implements distinct warm-up logic to validate connectivity before the real-time conversation begins. The responses-api backend sends minimal validation requests to /v1/responses, while the chat-completions implementation targets /v1/chat/completions via its warmup method (lines 94–102 in chat_completions_language_model.py).

When to Use Each Backend

Use responses-api when:

  • You require compatibility with older vLLM builds or providers that only implement the legacy Responses streaming specification
  • Your deployment scenario benefits from the simplest, unconverted Realtime format without message transformation overhead
  • You do not require the additional reasoning_effort control knob for managing model reasoning stages

Use chat-completions when:

  • You are running self-hosted vLLM, llama.cpp, or similar servers where tool-call streaming is more reliable via the Chat-Completions endpoint (particularly relevant for vLLM issue #312)
  • You need explicit control over reasoning effort via --responses_api_reasoning_effort to suppress thinking on providers that ignore standard disable flags
  • Your application requires strict adherence to modern OpenAI Chat-Completions multimodal formatting standards

CLI and Python Configuration Examples

Switching Backends via CLI

The backend selection is controlled via the --llm_backend flag, parsed in module_arguments.py and injected during pipeline construction in src/speech_to_speech/s2s_pipeline.py (lines 104–110). The default value is responses-api.

Using the default responses-api backend:

speech-to-speech \
    --mode realtime \
    --stt parakeet-tdt \
    --llm_backend responses-api \
    --tts qwen3 \
    --model_name "gpt-4o-mini" \
    --responses_api_api_key "$OPENAI_API_KEY" \
    --responses_api_stream

Switching to chat-completions with reasoning control:

speech-to-speech \
    --mode realtime \
    --stt parakeet-tdt \
    --llm_backend chat-completions \
    --tts qwen3 \
    --model_name "Qwen/Qwen3-4B-Instruct-2507" \
    --responses_api_base_url "http://localhost:8000/v1" \
    --responses_api_stream \
    --responses_api_reasoning_effort none

Direct Python Client Access

OpenAI client (responses-api):

from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:8000/v1",
    api_key="any-string",
)

response = client.responses.create(
    model="gpt-4o-mini",
    messages=[{"role": "user", "content": "Hello"}],
    stream=True,
)

OpenAI client (chat-completions with reasoning override):

from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:8000/v1",
    api_key="any-string",
)

response = client.chat.completions.create(
    model="gpt-4o-mini",
    messages=[{"role": "user", "content": "Hello"}],
    stream=True,
    extra_body={"reasoning_effort": "none"},
)

Summary

  • responses-api (/v1/responses) uses ResponsesApiModelHandler with raw Realtime format streaming, suitable for legacy compatibility and simple deployments
  • chat-completions (/v1/chat/completions) uses ChatCompletionsApiModelHandler with message conversion logic (_chat_messages) and stricter schema compliance, preferred for modern vLLM tool-calling
  • Both backends share base arguments from ResponsesApiLanguageModelHandlerArguments, but chat-completions uniquely adds --responses_api_reasoning_effort for fine-grained reasoning control
  • Selection occurs via --llm_backend flag, with pipeline instantiation handled in s2s_pipeline.py at lines 104–110
  • Internal event emission remains identical across both backends, ensuring seamless interoperability with the rest of the Speech-to-Speech pipeline

Frequently Asked Questions

What CLI flag selects between responses-api and chat-completions?

Use --llm_backend followed by either responses-api (default) or chat-completions. This flag is parsed in module_arguments.py and determines which handler class—ResponsesApiModelHandler or ChatCompletionsApiModelHandler—is instantiated during pipeline construction in s2s_pipeline.py.

Do both backends support tool calling in Speech-to-Speech?

Yes. Both handlers generate identical ToolCall internal events that the pipeline processes downstream. However, the chat-completions backend provides more reliable tool-call streaming for certain self-hosted providers like vLLM, while responses-api uses an older delta format that may encounter compatibility issues with specific model implementations.

Why does chat-completions require message conversion while responses-api does not?

The chat-completions backend must transform Realtime API message shapes into strict Chat-Completions schema compliance. Specifically, the _chat_messages method in chat_completions_language_model.py converts multimodal inputs from input_text/input_image to text/image_url formats and ensures tool arguments are JSON strings, whereas responses-api transmits the native Realtime format directly without transformation.

Can I use the same API key and base URL for both backends?

Yes. Both backends utilize identical CLI flags for connection parameters (--responses_api_base_url, --responses_api_api_key) inherited from ResponsesApiLanguageModelHandlerArguments. Simply change --llm_backend to switch protocols while keeping authentication configuration constant, though ensure the actual endpoint (/v1/responses vs /v1/chat/completions) is available at the specified base URL.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →