Responses API vs Chat Completions LLM Backends in HuggingFace Speech-to-Speech
The Responses-API backend targets the legacy /v1/responses endpoint with flat tool schemas and discrete event streaming, while the Chat-Completions backend uses the modern /v1/chat/completions endpoint with nested function tools and incremental delta streaming.
The HuggingFace speech-to-speech repository provides two concrete implementations of the BaseOpenAICompatibleHandler abstract class for integrating OpenAI-compatible language models. While both responses-api and chat-completions backends handle the same core functionality, they differ significantly in request serialization, tool handling, and streaming protocols. These distinctions determine which handler to deploy based on your specific LLM provider's API specifications.
Endpoint and Protocol Architecture
Both backends inherit from BaseOpenAICompatibleHandler but communicate with distinct HTTP endpoints.
Responses-API (ResponsesApiModelHandler) sends requests to POST /v1/responses and processes openai.types.responses events such as ResponseTextDeltaEvent, ResponseOutputItemDoneEvent, and ResponseFunctionToolCall.
Chat-Completions (ChatCompletionsApiModelHandler) sends requests to POST /v1/chat/completions and processes openai.Stream[ChatCompletionChunk] objects containing delta.tool_calls and delta.content fields.
These architectural differences require specific payload formatting and streaming logic for each backend.
Request Payload Serialization
The backends use different serialization paths to convert internal Chat objects into API-compatible payloads.
In src/speech_to_speech/LLM/responses_api_language_model.py, the Responses-API backend calls Chat.to_responses_api_chat() to generate a list of message objects with type="message" fields. Tools pass through unchanged because the format already matches the Responses API specification.
In src/speech_to_speech/LLM/chat_completions_language_model.py, the Chat-Completions backend first calls Chat.to_transformers_chat(), then applies _chat_messages() to nest function tools under a function key and convert Realtime-style content (input_text, input_image) into Chat-Completions shapes (text, image_url).
# Responses-API serialization
payload_responses = ResponsesApiModelHandler()._serialize(chat)
# Chat-Completions serialization (includes conversion)
payload_chat = ChatCompletionsApiModelHandler()._serialize(chat)
Tool Definition Conversion
Tool handling represents a major divergence between the two backends.
Responses-API accepts flat tool definitions directly. The format requires no conversion because the handler passes tools unchanged to client.responses.create().
Chat-Completions requires tool restructuring. The _to_chat_tools() method wraps flat {"type":"function",...} definitions into ChatCompletionToolParam objects with a nested function key:
from speech_to_speech.LLM.chat_completions_language_model import _to_chat_tools
flat_tools = [
{"type": "function", "name": "search", "description": "Search the web", "parameters": {"type": "object"}}
]
chat_tools = _to_chat_tools(flat_tools)
# Result: [{'type': 'function', 'function': {'name': 'search', ...}}]
Additionally, the Chat-Completions backend converts tool_choice parameters using _to_chat_tool_choice(), while the Responses-API backend preserves the original tool choice format.
Streaming Protocol Differences
Streaming implementations differ in event types and tool-call accumulation strategies.
Responses-API yields discrete events from the stream. When a ResponseFunctionToolCall appears, the handler immediately emits a complete ToolCall event with a regenerated ID via _generate_id().
Chat-Completions accumulates partial tool-call fragments across multiple deltas. The _iter_stream_events() method stitches together incremental pieces using an index-keyed dictionary before yielding a complete ToolCall after the stream ends.
Refusal handling also varies:
- Responses-API: Refusals appear as
ResponseOutputMessagewith arefusalfield, converted to anAssistantMessagecontainingAssistantContent(type="output_text"). - Chat-Completions: Refusals stream as
delta.refusal, emitted immediately asTextDeltaobjects wrapped inAssistantMessageinstances.
Multimodal Content Handling
The backends process multimodal inputs differently.
Responses-API directly forwards Realtime-style input_text and input_image parts without transformation.
Chat-Completions converts these parts via _to_chat_content_part(), transforming input_text into {"type": "text", ...} and input_image into {"type": "image_url", ...} objects suitable for the Chat-Completions schema.
Instantiation Examples
Configure each backend by instantiating the appropriate handler class with your model parameters.
Responses-API Backend
from speech_to_speech.LLM.responses_api_language_model import ResponsesApiModelHandler
import threading, queue
handler = ResponsesApiModelHandler(
threading.Event(),
queue.Queue(),
queue.Queue(),
setup_kwargs=dict(
model_name="meta-llama/Meta-Llama-3-8B-Instruct",
base_url="http://my-server/v1",
api_key="my-key",
stream=True,
disable_thinking=False,
),
)
Chat-Completions Backend
from speech_to_speech.LLM.chat_completions_language_model import ChatCompletionsApiModelHandler
import threading, queue
handler = ChatCompletionsApiModelHandler(
threading.Event(),
queue.Queue(),
queue.Queue(),
setup_kwargs=dict(
model_name="meta-llama/Meta-Llama-3-8B-Instruct",
base_url="http://my-server/v1",
api_key="my-key",
stream=True,
disable_thinking=False,
),
)
Extra Body and Provider-Specific Logic
The Chat-Completions backend includes conditional logic for extra body parameters. According to the implementation in BaseOpenAICompatibleHandler._build_extra_body(), the handler sends extra_body (such as {"chat_template_kwargs":{"enable_thinking":False}}) only when the base URL is not OpenAI's official endpoint. The Responses-API backend always transmits its self._extra_body configuration regardless of the endpoint.
Summary
- Responses-API (
src/speech_to_speech/LLM/responses_api_language_model.py) targets/v1/responseswith flat tool schemas, discrete event streaming, and direct multimodal forwarding. - Chat-Completions (
src/speech_to_speech/LLM/chat_completions_language_model.py) targets/v1/chat/completionswith nested tool definitions, incremental delta streaming, and Realtime-to-Chat conversion logic. - Both implement
BaseOpenAICompatibleHandlerbut differ in serialization methods (to_responses_api_chat()vsto_transformers_chat()plus_chat_messages()). - Tool calls arrive as complete events in Responses-API but require accumulation from deltas in Chat-Completions.
- Choose Responses-API for legacy OpenAI-compatible endpoints; use Chat-Completions for modern providers following the standard OpenAI schema.
Frequently Asked Questions
Which backend should I use for OpenAI's official API?
Use the Chat-Completions backend for OpenAI's official API. The Responses-API endpoint represents a legacy protocol, while the Chat-Completions endpoint (/v1/chat/completions) is the current standard. The Chat-Completions handler also automatically omits extra body parameters when connecting to OpenAI's official domain, ensuring compatibility with OpenAI's strict schema validation.
How does tool calling differ between the two backends?
The Responses-API backend receives complete tool calls as ResponseFunctionToolCall events and yields them immediately with regenerated IDs. The Chat-Completions backend receives fragmented tool call deltas across multiple chunks and must accumulate them using an index-keyed dictionary in _iter_stream_events() before emitting a complete ToolCall. Additionally, Chat-Completions requires wrapping flat tool definitions into nested function objects via _to_chat_tools().
Can I switch between backends without changing my chat history format?
Yes. Both backends accept the same internal Chat object from src/speech_to_speech/LLM/chat.py. The handlers automatically convert the chat history to their respective API formats using to_responses_api_chat() or to_transformers_chat() followed by _chat_messages(). Your application code can instantiate either ResponsesApiModelHandler or ChatCompletionsApiModelHandler using the same chat history without manual conversion.
What happens to multimodal content in each backend?
The Responses-API backend forwards Realtime-style input_text and input_image parts directly without transformation. The Chat-Completions backend converts these parts through _to_chat_content_part(), translating input_text to {"type": "text", ...} and input_image to {"type": "image_url", ...} to match the Chat-Completions schema requirements.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →