Differences Between responses-api and chat-completions LLM Backends in Speech-to-Speech

The responses-api backend communicates via the legacy /v1/responses endpoint with basic streaming support, while the chat-completions backend uses the modern /v1/chat/completions protocol featuring robust tool-call streaming and extended reasoning controls.

The Hugging Face speech-to-speech (STS) repository provides two interchangeable OpenAI-compatible HTTP endpoints for language model inference. Understanding the architectural differences between the responses-api and chat-completions LLM backends allows you to optimize your voice-agent pipeline for specific deployment scenarios, from legacy server compatibility to advanced tool-calling with self-hosted vLLM instances.

Core Architectural Differences

Endpoint Specifications and Handler Classes

The two backends implement distinct handler classes that target different API endpoints. The responses-api backend uses ResponsesApiModelHandler located in src/speech_to_speech/LLM/responses_api_language_model.py, which defaults to sending requests to POST /v1/responses. Conversely, the chat-completions backend implements ChatCompletionsApiModelHandler in src/speech_to_speech/LLM/chat_completions_language_model.py, targeting the POST /v1/chat/completions endpoint.

Both handlers share a common argument foundation through ResponsesApiLanguageModelHandlerArguments defined in src/speech_to_speech/arguments_classes/responses_api_language_model_arguments.py. This base class provides CLI flags such as --responses_api_base_url, --responses_api_api_key, and --responses_api_stream that work identically across both backends.

Message Conversion and Payload Adaptation

The chat-completions backend performs critical message transformation through the _chat_messages method (lines 55-66 in chat_completions_language_model.py). This conversion logic performs two essential adaptations:

  • Tool Call Serialization: Ensures tool_calls.arguments are properly formatted as JSON strings rather than raw objects
  • Multimodal Rewriting: Transforms Realtime API format elements (input_text, input_image) into Chat-Completions compatible shapes (text, image_url)

The responses-api backend sends chat payloads as-is without transformation, maintaining the native OpenAI Realtime format.

Streaming Protocol Implementation

Both backends implement an _iter_stream_events method to process streaming responses, but handle different response schemas:

  • responses-api: Consumes Stream[ResponseChunk] objects and extracts tool-call deltas from the Responses-specific format
  • chat-completions: Processes Stream[ChatCompletionChunk] objects using the stricter Chat-Completions schema, which provides more reliable delta parsing for complex tool invocations

Despite these differences, both handlers ultimately emit identical internal events—including AssistantMessage, ToolCall, TextDelta, and Usage—ensuring the rest of the STS pipeline (VAD → STT → TTS) remains backend-agnostic.

Feature Comparison: Tool Calling and Reasoning Control

Tool-Call Reliability and Format

The responses-api backend uses ResponseFunctionToolCall objects with streaming deltas arriving in choices[].delta.tool_calls. While functional, this format occasionally exhibits compatibility issues with certain server implementations.

The chat-completions backend leverages the mature Chat-Completions streaming protocol, which has proven more reliable for vLLM and Qwen tool-calling scenarios (see repository issue #312). This protocol handles nested tool-call arguments more robustly during streaming inference.

Reasoning and Thinking Parameters

Both backends respect the responses_api_disable_thinking flag to control model reasoning behavior. However, chat-completions extends this capability through an additional parameter defined in ChatCompletionsLanguageModelHandlerArguments at src/speech_to_speech/arguments_classes/chat_completions_language_model_arguments.py.

The responses_api_reasoning_effort parameter allows you to send extra_body={'reasoning_effort': <value>} to force providers to skip or reduce reasoning effort when they ignore the generic disable_thinking flag. This proves essential when deploying models through certain vLLM configurations that require explicit reasoning controls in the request body.

Warm-up and Connection Handling

Each backend implements a warmup method to validate connectivity before processing real-time audio:

  • responses-api: Sends a minimal request to /v1/responses
  • chat-completions: Sends a minimal request to /v1/chat/completions (see lines 94-102 in chat_completions_language_model.py)

These warm-up routines ensure the endpoint is responsive before the voice conversation begins, preventing cold-start latency during active sessions.

When to Use Each Backend

Select responses-api when you require broad compatibility with older OpenAI-compatible providers or prefer the simplest streaming format without message conversion overhead. This backend works reliably with any server implementing the legacy Responses specification.

Choose chat-completions for the following scenarios:

  • Self-hosted vLLM or llama.cpp servers where tool-call streaming proves more stable through the Chat-Completions endpoint rather than the Responses API
  • Advanced reasoning control when you need the responses_api_reasoning_effort knob to suppress thinking on providers that bypass the standard disable_thinking flag
  • Production tool-calling with complex multimodal inputs requiring the robust payload adaptation provided by the _chat_messages conversion layer

Configuration and Code Examples

CLI Configuration

Activate the responses-api backend (default) using standard flags:

speech-to-speech \
    --mode realtime \
    --stt parakeet-tdt \
    --llm_backend responses-api \
    --tts qwen3 \
    --model_name "gpt-4o-mini" \
    --responses_api_api_key "$OPENAI_API_KEY" \
    --responses_api_stream

Switch to the chat-completions backend with extended reasoning controls:

speech-to-speech \
    --mode realtime \
    --stt parakeet-tdt \
    --llm_backend chat-completions \
    --tts qwen3 \
    --model_name "Qwen/Qwen3-4B-Instruct-2507" \
    --responses_api_base_url "http://localhost:8000/v1" \
    --responses_api_stream \
    --responses_api_reasoning_effort none

Python Client Integration

Direct OpenAI client usage with the responses-api endpoint:

from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:8000/v1",
    api_key="any-string",
)

response = client.responses.create(
    model="gpt-4o-mini",
    messages=[{"role": "user", "content": "Hello"}],
    stream=True,
)

Direct OpenAI client usage with the chat-completions endpoint and reasoning control:

from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:8000/v1",
    api_key="any-string",
)

response = client.chat.completions.create(
    model="gpt-4o-mini",
    messages=[{"role": "user", "content": "Hello"}],
    stream=True,
    extra_body={"reasoning_effort": "none"},
)

Pipeline Selection

The backend selection occurs during pipeline construction in src/speech_to_speech/s2s_pipeline.py (lines 104-110), where the --llm_backend flag value determines which handler class instantiates the LLM module. Both handlers accept identical base configuration parameters, simplifying backend swaps without modifying downstream audio processing components.

Summary

  • Endpoint Differences: responses-api targets /v1/responses while chat-completions uses /v1/chat/completions
  • Handler Classes: ResponsesApiModelHandler vs ChatCompletionsApiModelHandler in their respective files under src/speech_to_speech/LLM/
  • Message Processing: chat-completions converts payloads via _chat_messages to ensure JSON string arguments and proper multimodal formatting
  • Tool Calling: chat-completions provides more reliable streaming for vLLM and Qwen tool invocations
  • Reasoning Control: chat-completions adds responses_api_reasoning_effort via ChatCompletionsLanguageModelHandlerArguments for providers ignoring standard disable flags
  • Pipeline Compatibility: Both backends emit identical internal events, ensuring seamless integration with STT and TTS components

Frequently Asked Questions

What is the default LLM backend in speech-to-speech?

The responses-api backend serves as the default when you omit the --llm_backend flag. This default is enforced during argument parsing in the pipeline construction phase, ensuring backward compatibility with existing deployments that expect the legacy /v1/responses endpoint behavior.

Can I switch between backends without modifying my pipeline code?

Yes. Both backends emit identical internal event types (AssistantMessage, ToolCall, Usage), making the rest of the pipeline agnostic to your choice. You only need to change the --llm_backend flag and potentially adjust the --responses_api_reasoning_effort parameter when switching to chat-completions.

Why does chat-completions have an extra reasoning_effort parameter?

The responses_api_reasoning_effort parameter exists in ChatCompletionsLanguageModelHandlerArguments to address providers (such as certain vLLM versions) that ignore the generic disable_thinking flag. It sends reasoning controls via extra_body, forcing the model to skip or reduce reasoning effort when standard flags fail.

Which backend should I use with self-hosted vLLM servers?

Use chat-completions when running self-hosted vLLM or llama.cpp instances, particularly if you rely on tool-calling functionality. The Chat-Completions streaming protocol handles tool-call deltas more reliably than the Responses API format, resolving compatibility issues documented in repository issue #312.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →