# Responses API vs Chat Completions LLM Backends in HuggingFace Speech-to-Speech

> Compare Responses API vs Chat Completions LLM backends in HuggingFace Speech-to-Speech. Understand differences in tool schemas and streaming for efficient integration.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: comparisons
- Published: 2026-07-11

---

**The Responses-API backend targets the legacy `/v1/responses` endpoint with flat tool schemas and discrete event streaming, while the Chat-Completions backend uses the modern `/v1/chat/completions` endpoint with nested function tools and incremental delta streaming.**

The HuggingFace *speech-to-speech* repository provides two concrete implementations of the `BaseOpenAICompatibleHandler` abstract class for integrating OpenAI-compatible language models. While both **responses-api** and **chat-completions** backends handle the same core functionality, they differ significantly in request serialization, tool handling, and streaming protocols. These distinctions determine which handler to deploy based on your specific LLM provider's API specifications.

## Endpoint and Protocol Architecture

Both backends inherit from `BaseOpenAICompatibleHandler` but communicate with distinct HTTP endpoints.

**Responses-API** (`ResponsesApiModelHandler`) sends requests to `POST /v1/responses` and processes `openai.types.responses` events such as `ResponseTextDeltaEvent`, `ResponseOutputItemDoneEvent`, and `ResponseFunctionToolCall`.

**Chat-Completions** (`ChatCompletionsApiModelHandler`) sends requests to `POST /v1/chat/completions` and processes `openai.Stream[ChatCompletionChunk]` objects containing `delta.tool_calls` and `delta.content` fields.

These architectural differences require specific payload formatting and streaming logic for each backend.

## Request Payload Serialization

The backends use different serialization paths to convert internal `Chat` objects into API-compatible payloads.

In [`src/speech_to_speech/LLM/responses_api_language_model.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/responses_api_language_model.py), the Responses-API backend calls `Chat.to_responses_api_chat()` to generate a list of message objects with `type="message"` fields. Tools pass through unchanged because the format already matches the Responses API specification.

In [`src/speech_to_speech/LLM/chat_completions_language_model.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/chat_completions_language_model.py), the Chat-Completions backend first calls `Chat.to_transformers_chat()`, then applies `_chat_messages()` to nest function tools under a `function` key and convert Realtime-style content (`input_text`, `input_image`) into Chat-Completions shapes (`text`, `image_url`).

```python

# Responses-API serialization

payload_responses = ResponsesApiModelHandler()._serialize(chat)

# Chat-Completions serialization (includes conversion)

payload_chat = ChatCompletionsApiModelHandler()._serialize(chat)

```

## Tool Definition Conversion

Tool handling represents a major divergence between the two backends.

**Responses-API** accepts flat tool definitions directly. The format requires no conversion because the handler passes tools unchanged to `client.responses.create()`.

**Chat-Completions** requires tool restructuring. The `_to_chat_tools()` method wraps flat `{"type":"function",...}` definitions into `ChatCompletionToolParam` objects with a nested `function` key:

```python
from speech_to_speech.LLM.chat_completions_language_model import _to_chat_tools

flat_tools = [
    {"type": "function", "name": "search", "description": "Search the web", "parameters": {"type": "object"}}
]

chat_tools = _to_chat_tools(flat_tools)

# Result: [{'type': 'function', 'function': {'name': 'search', ...}}]

```

Additionally, the Chat-Completions backend converts `tool_choice` parameters using `_to_chat_tool_choice()`, while the Responses-API backend preserves the original tool choice format.

## Streaming Protocol Differences

Streaming implementations differ in event types and tool-call accumulation strategies.

**Responses-API** yields discrete events from the stream. When a `ResponseFunctionToolCall` appears, the handler immediately emits a complete `ToolCall` event with a regenerated ID via `_generate_id()`.

**Chat-Completions** accumulates partial tool-call fragments across multiple deltas. The `_iter_stream_events()` method stitches together incremental pieces using an index-keyed dictionary before yielding a complete `ToolCall` after the stream ends.

Refusal handling also varies:
- **Responses-API**: Refusals appear as `ResponseOutputMessage` with a `refusal` field, converted to an `AssistantMessage` containing `AssistantContent(type="output_text")`.
- **Chat-Completions**: Refusals stream as `delta.refusal`, emitted immediately as `TextDelta` objects wrapped in `AssistantMessage` instances.

## Multimodal Content Handling

The backends process multimodal inputs differently.

**Responses-API** directly forwards Realtime-style `input_text` and `input_image` parts without transformation.

**Chat-Completions** converts these parts via `_to_chat_content_part()`, transforming `input_text` into `{"type": "text", ...}` and `input_image` into `{"type": "image_url", ...}` objects suitable for the Chat-Completions schema.

## Instantiation Examples

Configure each backend by instantiating the appropriate handler class with your model parameters.

### Responses-API Backend

```python
from speech_to_speech.LLM.responses_api_language_model import ResponsesApiModelHandler
import threading, queue

handler = ResponsesApiModelHandler(
    threading.Event(),
    queue.Queue(),
    queue.Queue(),
    setup_kwargs=dict(
        model_name="meta-llama/Meta-Llama-3-8B-Instruct",
        base_url="http://my-server/v1",
        api_key="my-key",
        stream=True,
        disable_thinking=False,
    ),
)

```

### Chat-Completions Backend

```python
from speech_to_speech.LLM.chat_completions_language_model import ChatCompletionsApiModelHandler
import threading, queue

handler = ChatCompletionsApiModelHandler(
    threading.Event(),
    queue.Queue(),
    queue.Queue(),
    setup_kwargs=dict(
        model_name="meta-llama/Meta-Llama-3-8B-Instruct",
        base_url="http://my-server/v1",
        api_key="my-key",
        stream=True,
        disable_thinking=False,
    ),
)

```

## Extra Body and Provider-Specific Logic

The Chat-Completions backend includes conditional logic for extra body parameters. According to the implementation in `BaseOpenAICompatibleHandler._build_extra_body()`, the handler sends `extra_body` (such as `{"chat_template_kwargs":{"enable_thinking":False}}`) only when the base URL is not OpenAI's official endpoint. The Responses-API backend always transmits its `self._extra_body` configuration regardless of the endpoint.

## Summary

- **Responses-API** ([`src/speech_to_speech/LLM/responses_api_language_model.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/responses_api_language_model.py)) targets `/v1/responses` with flat tool schemas, discrete event streaming, and direct multimodal forwarding.
- **Chat-Completions** ([`src/speech_to_speech/LLM/chat_completions_language_model.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/chat_completions_language_model.py)) targets `/v1/chat/completions` with nested tool definitions, incremental delta streaming, and Realtime-to-Chat conversion logic.
- Both implement `BaseOpenAICompatibleHandler` but differ in serialization methods (`to_responses_api_chat()` vs `to_transformers_chat()` plus `_chat_messages()`).
- Tool calls arrive as complete events in Responses-API but require accumulation from deltas in Chat-Completions.
- Choose Responses-API for legacy OpenAI-compatible endpoints; use Chat-Completions for modern providers following the standard OpenAI schema.

## Frequently Asked Questions

### Which backend should I use for OpenAI's official API?

Use the **Chat-Completions** backend for OpenAI's official API. The Responses-API endpoint represents a legacy protocol, while the Chat-Completions endpoint (`/v1/chat/completions`) is the current standard. The Chat-Completions handler also automatically omits extra body parameters when connecting to OpenAI's official domain, ensuring compatibility with OpenAI's strict schema validation.

### How does tool calling differ between the two backends?

The **Responses-API** backend receives complete tool calls as `ResponseFunctionToolCall` events and yields them immediately with regenerated IDs. The **Chat-Completions** backend receives fragmented tool call deltas across multiple chunks and must accumulate them using an index-keyed dictionary in `_iter_stream_events()` before emitting a complete `ToolCall`. Additionally, Chat-Completions requires wrapping flat tool definitions into nested `function` objects via `_to_chat_tools()`.

### Can I switch between backends without changing my chat history format?

Yes. Both backends accept the same internal `Chat` object from [`src/speech_to_speech/LLM/chat.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/chat.py). The handlers automatically convert the chat history to their respective API formats using `to_responses_api_chat()` or `to_transformers_chat()` followed by `_chat_messages()`. Your application code can instantiate either `ResponsesApiModelHandler` or `ChatCompletionsApiModelHandler` using the same chat history without manual conversion.

### What happens to multimodal content in each backend?

The **Responses-API** backend forwards Realtime-style `input_text` and `input_image` parts directly without transformation. The **Chat-Completions** backend converts these parts through `_to_chat_content_part()`, translating `input_text` to `{"type": "text", ...}` and `input_image` to `{"type": "image_url", ...}` to match the Chat-Completions schema requirements.