Difference Between responses-api and chat-completions LLM Backends in Speech-to-Speech

The responses-api backend connects to OpenAI's legacy /v1/responses endpoint, while chat-completions uses the modern /v1/chat/completions API with enhanced tool-call streaming and message format conversion.

The Hugging Face Speech-to-Speech (STS) pipeline supports two interchangeable OpenAI-compatible HTTP backends for Large Language Model (LLM) inference. Understanding the difference between responses-api and chat-completions LLM backends helps you select the right protocol for your deployment, whether you need legacy compatibility with older providers or robust tool-calling support with self-hosted servers.

What Are the LLM Backends?

Speech-to-Speech can communicate with language models through two distinct HTTP endpoints that share a common argument foundation. Both backends inherit from ResponsesApiLanguageModelHandlerArguments defined in src/speech_to_speech/arguments_classes/responses_api_language_model_arguments.py, which provides CLI flags like --responses_api_base_url, --responses_api_api_key, and --responses_api_stream.

The chat-completions backend extends these arguments through ChatCompletionsLanguageModelHandlerArguments in src/speech_to_speech/arguments_classes/chat_completions_language_model_arguments.py, adding the unique responses_api_reasoning_effort parameter for providers that require explicit reasoning control.

Key Architectural Differences

Endpoint and Handler Classes

Each backend uses a distinct endpoint and handler implementation:

Message Format Conversion

The Chat-Completions backend performs explicit message adaptation before sending requests. In src/speech_to_speech/LLM/chat_completions_language_model.py (lines 55-66), the _chat_messages method:

  • Ensures tool_calls.arguments are serialized as JSON strings rather than raw objects.
  • Rewrites multimodal content from Realtime API shapes (input_text, input_image) to Chat-Completions format (text, image_url).

The Responses-API backend sends the chat payload as-is in OpenAI Realtime format without conversion.

Tool-Call Streaming

Both handlers implement _iter_stream_events to process streaming responses, but they handle different response schemas:

  • Responses-API: Gathers tool-call deltas from Stream[ResponseChunk] objects using the older ResponseFunctionToolCall format.
  • Chat-Completions: Processes Stream[ChatCompletionChunk] objects (lines 5-12 in chat_completions_language_model.py), emitting AssistantMessage, TextDelta, and ToolCall events using the stricter Chat-Completions schema that is more reliable for vLLM and Qwen tool-calling.

Reasoning Control

The backends differ in how they handle reasoning or "thinking" modes:

  • Responses-API: Controlled solely by the responses_api_disable_thinking flag.
  • Chat-Completions: Supports both responses_api_disable_thinking and the additional responses_api_reasoning_effort parameter, which sends extra_body={'reasoning_effort': <value>} to force providers to skip or reduce reasoning effort when the generic flag is ignored.

When to Use Each Backend

Use responses-api when you need compatibility with older OpenAI-compatible providers or simple deployments where the legacy streaming format is sufficient.

Use chat-completions when:

  • Running self-hosted vLLM or llama.cpp servers where the Chat-Completions endpoint provides more reliable tool-call streaming (see issue #312).
  • You need the responses_api_reasoning_effort knob to suppress reasoning on providers that ignore the standard disable_thinking flag.
  • Working with multimodal inputs that require conversion from Realtime API formats to standard Chat-Completions shapes.

Configuration and CLI Examples

Select your backend using the --llm_backend flag parsed in module_arguments.py and injected into the pipeline in src/speech_to_speech/s2s_pipeline.py (lines 104-110).

Using the Responses-API backend (default):

speech-to-speech \
    --mode realtime \
    --stt parakeet-tdt \
    --llm_backend responses-api \
    --tts qwen3 \
    --model_name "gpt-4o-mini" \
    --responses_api_api_key "$OPENAI_API_KEY" \
    --responses_api_stream

Switching to the Chat-Completions backend:

speech-to-speech \
    --mode realtime \
    --stt parakeet-tdt \
    --llm_backend chat-completions \
    --tts qwen3 \
    --model_name "Qwen/Qwen3-4B-Instruct-2507" \
    --responses_api_base_url "http://localhost:8000/v1" \
    --responses_api_stream \
    --responses_api_reasoning_effort none

Direct OpenAI client usage (Responses-API):

from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:8000/v1",
    api_key="any-string",
)

response = client.responses.create(
    model="gpt-4o-mini",
    messages=[{"role": "user", "content": "Hello"}],
    stream=True,
)

Direct OpenAI client usage (Chat-Completions):

from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:8000/v1",
    api_key="any-string",
)

response = client.chat.completions.create(
    model="gpt-4o-mini",
    messages=[{"role": "user", "content": "Hello"}],
    stream=True,
    extra_body={"reasoning_effort": "none"},
)

Summary

  • Both backends inherit from ResponsesApiLanguageModelHandlerArguments and share base CLI flags, but chat-completions adds the responses_api_reasoning_effort parameter.
  • Responses-API uses /v1/responses with ResponsesApiModelHandler and sends messages in Realtime format without conversion.
  • Chat-Completions uses /v1/chat/completions with ChatCompletionsApiModelHandler, converting tool arguments to JSON strings and rewriting multimodal content.
  • The Chat-Completions backend provides more reliable tool-call streaming for vLLM and Qwen deployments.
  • Select backends via --llm_backend (default: responses-api) in s2s_pipeline.py.

Frequently Asked Questions

Which backend should I use with vLLM?

Use the chat-completions backend when running vLLM servers. According to the source code in chat_completions_language_model.py, this backend provides more reliable tool-call streaming compared to the Responses-API format, which can be flaky with certain vLLM configurations (see issue #312).

How do I disable reasoning or "thinking" in the LLM?

For the Responses-API backend, use the --responses_api_disable_thinking flag. For the Chat-Completions backend, you can use both --responses_api_disable_thinking and --responses_api_reasoning_effort none. The latter sends extra_body={'reasoning_effort': 'none'} to providers that ignore the standard disable flag.

Are the backends interchangeable in the pipeline?

Yes. Both handlers ultimately produce the same internal events (AssistantMessage, ToolCall, Usage, etc.), making the rest of the pipeline (VAD → STT → TTS) agnostic to which backend you choose. The selection only affects the HTTP protocol and message formatting layer.

What is the default LLM backend?

The default is responses-api. You can override this by setting --llm_backend chat-completions in your CLI arguments, which the pipeline constructor reads from module_arguments.py and applies in s2s_pipeline.py during handler initialization.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →