# How to Implement Response Streaming with Responses API Backend in Speech-to-Speech

> Implement response streaming with the Responses API backend. Enable real-time token delivery in speech-to-speech for incremental TextDelta events. Learn how to set responses_api_stream=True.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: how-to-guide
- Published: 2026-07-30

---

**Enable real-time token delivery by setting `responses_api_stream=True` in `ResponsesApiLanguageModelHandlerArguments`, which propagates the stream flag through the OpenAI-compatible client and yields incremental `TextDelta` events as the provider generates them.**

The huggingface/speech-to-speech repository provides a production-ready pipeline for real-time voice conversations, supporting OpenAI-compatible Responses API backends from providers like Together, OpenAI, and Azure. Understanding how to implement response streaming with the Responses API backend is essential for building low-latency applications that display text as it generates rather than waiting for complete responses. This guide examines the three-layer architecture that controls streaming behavior and provides executable code samples for both high-level pipeline integration and low-level handler manipulation.

## Understanding the Streaming Architecture

Response streaming operates through three coordinated layers in the codebase, each handling a specific aspect of the data flow from configuration to event consumption.

### Configuration Layer

The `responses_api_stream` parameter in `ResponsesApiLanguageModelHandlerArguments` serves as the master switch for streaming behavior. Located in [`src/speech_to_speech/arguments_classes/responses_api_language_model_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/responses_api_language_model_arguments.py) at lines 21-27, this dataclass field defaults to `True`:

```python

# src/speech_to_speech/arguments_classes/responses_api_language_model_arguments.py

responses_api_stream: bool = field(
    default=True,
    metadata={"help": "Enable continuous token‑by‑token delivery via the Responses API."},
)

```

When you instantiate the pipeline with `--language-model-type responses-api`, the [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py) file passes this value to the handler constructor at lines 855-861, ensuring the boolean propagates through the entire request lifecycle.

### Request Layer

The `ResponsesApiModelHandler._request` method forwards the stream flag directly to the underlying OpenAI client. In [`src/speech_to_speech/LLM/responses_api_language_model.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/responses_api_language_model.py) at lines 97-104, the method signature binds the instance attribute `self.stream` to the API call:

```python

# src/speech_to_speech/LLM/responses_api_language_model.py

def _request(self, api_input, optional_kwargs):
    return self.client.responses.create(
        model=self.model_name,
        input=api_input,
        stream=self.stream,
        extra_body=self._extra_body,
        timeout=self.request_timeout,
        **optional_kwargs,
    )

```

Setting `stream=True` causes the client to return an `openai.Stream` object rather than a blocking response object, establishing the HTTP connection for server-sent events.

### Event Iteration Layer

When streaming is enabled, the handler processes the raw stream through `_iter_stream_events`, converting low-level SDK events into high-level `ProviderEvent` objects. This method in [`src/speech_to_speech/LLM/responses_api_language_model.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/responses_api_language_model.py) (lines 13-30) yields `TextDelta` instances containing incremental text fragments:

```python

# src/speech_to_speech/LLM/responses_api_language_model.py

def _iter_stream_events(self, api_response: Stream) -> Iterator[ProviderEvent]:
    for raw_event in api_response:
        if isinstance(raw_event, ResponseTextDeltaEvent):
            yield TextDelta(text=raw_event.delta)
        elif isinstance(raw_event, ResponseOutputItemDoneEvent):
            ...

```

The pipeline consumes these events asynchronously, enabling real-time display of generated content without buffering the entire response.

## Configuring Response Streaming

You can enable streaming through command-line arguments or programmatic configuration depending on your deployment requirements.

**Command-line activation:**

```bash
python -m speech_to_speech \
  --language-model-type responses-api \
  --responses_api_stream True \
  --responses_api_language_model_name "gpt-5.4-mini"

```

**Programmatic configuration:**

```python
from speech_to_speech.s2s_pipeline import SpeechToSpeechPipeline
from speech_to_speech.arguments_classes.responses_api_language_model_arguments import (
    ResponsesApiLanguageModelHandlerArguments,
)

args = ResponsesApiLanguageModelHandlerArguments(
    model_name="gpt-5.4-mini",
    responses_api_stream=True,
)

pipeline = SpeechToSpeechPipeline(
    language_model_type="responses-api",
    responses_api_language_model_handler_kwargs=args,
)

```

## Practical Implementation Examples

### High-Level Pipeline Usage

This minimal script demonstrates streaming responses using the high-level pipeline interface:

```python

# demo_stream_responses.py

import argparse
from speech_to_speech.s2s_pipeline import SpeechToSpeechPipeline
from speech_to_speech.arguments_classes.responses_api_language_model_arguments import (
    ResponsesApiLanguageModelHandlerArguments,
)

parser = argparse.ArgumentParser()
parser.add_argument("--prompt", default="Tell me a short joke.")
args = parser.parse_args()

pipeline = SpeechToSpeechPipeline(
    language_model_type="responses-api",
    responses_api_language_model_handler_kwargs=ResponsesApiLanguageModelHandlerArguments(
        model_name="gpt-5.4-mini",
        responses_api_stream=True,
        responses_api_disable_thinking=False,
    ),
)

for event in pipeline.run(prompt=args.prompt):
    if isinstance(event, pipeline.provider_events.TextDelta):
        print(event.text, end="", flush=True)

```

Running this script prints tokens as they arrive from the provider, achieving true real-time streaming behavior suitable for live chat interfaces.

### Direct Handler Control

For fine-grained control over the streaming process, interact directly with `ResponsesApiModelHandler`:

```python

# direct_handler_stream.py

from speech_to_speech.LLM.responses_api_language_model import ResponsesApiModelHandler
from speech_to_speech.arguments_classes.responses_api_language_model_arguments import (
    ResponsesApiLanguageModelHandlerArguments,
)

args = ResponsesApiLanguageModelHandlerArguments(
    model_name="gpt-5.4-mini",
    responses_api_stream=True,
    responses_api_disable_thinking=True,
)

handler = ResponsesApiModelHandler(
    model_name=args.model_name,
    stream=args.responses_api_stream,
    disable_thinking=args.responses_api_disable_thinking,
    client=None,  # Uses OPENAI_API_KEY and BASE_URL env vars

)

api_input = [
    {"type": "message", "role": "system", "content": [{"type": "input_text", "text": "You are a helpful assistant"}]},
    {"type": "message", "role": "user", "content": [{"type": "input_text", "text": "What's the weather like today?"}]},
]

stream = handler._request(api_input, optional_kwargs={})
for ev in handler._iter_stream_events(stream):
    if isinstance(ev, handler.TextDelta):
        print(ev.text, end="", flush=True)
    elif isinstance(ev, handler.ToolCall):
        pass  # Handle function calls dynamically

```

This approach exposes the raw `openai.Stream` and allows custom processing of `ResponseTextDeltaEvent` objects before they reach the pipeline.

### Disabling Provider-Side Reasoning

Some providers emit "thinking" tokens alongside content. Disable this to stream only final output tokens:

```python
args = ResponsesApiLanguageModelHandlerArguments(
    model_name="gpt-5.4-mini",
    responses_api_stream=True,
    responses_api_disable_thinking=True,  # Disables chain-of-thought

)

```

When `responses_api_disable_thinking` is set to `True`, the handler injects `chat_template_kwargs={"enable_thinking": False}` into the request payload via the logic in [`src/speech_to_speech/LLM/responses_api_language_model.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/responses_api_language_model.py) lines 86-94.

## Summary

- **Three-layer control**: Streaming configuration flows from `responses_api_stream` in the arguments dataclass, through the `_request` method's `stream` parameter, to the `_iter_stream_events` generator that yields `TextDelta` objects.
- **Real-time delivery**: The `ResponsesApiModelHandler` converts OpenAI `Stream` objects into consumable events, printing tokens immediately as the LLM generates them.
- **Thinking suppression**: Use `responses_api_disable_thinking` to eliminate reasoning tokens and reduce bandwidth when only final output is required.
- **Flexible integration**: The pipeline supports both high-level `SpeechToSpeechPipeline` usage for standard workflows and direct `ResponsesApiModelHandler` manipulation for custom event processing.

## Frequently Asked Questions

### Which providers support the Responses API streaming mode?

The implementation works with any OpenAI-compatible endpoint, including Together AI, OpenAI, and Azure OpenAI Service. As long as the provider supports the Responses API schema and streaming parameters, the `ResponsesApiModelHandler` will correctly yield `TextDelta` events according to the huggingface/speech-to-speech source code.

### How do I disable streaming to receive complete responses?

Set `responses_api_stream=False` in your `ResponsesApiLanguageModelHandlerArguments` configuration. When disabled, the `_request` method passes `stream=False` to the client, causing `responses.create()` to return a complete response object rather than a `Stream`, and the handler skips the `_iter_stream_events` generator in favor of synchronous result processing.

### What is the difference between `TextDelta` and `ResponseOutputItemDoneEvent`?

`TextDelta` represents incremental content fragments delivered via `ResponseTextDeltaEvent` during active streaming, suitable for displaying partial text immediately. `ResponseOutputItemDoneEvent` signals the completion of an output item, indicating that the provider has finished generating that specific content block and no further deltas will arrive for that item.

### Can I use custom authentication headers with the streaming handler?

Yes. When initializing `ResponsesApiModelHandler` with `client=None`, the handler uses the default OpenAI client initialization which respects `OPENAI_API_KEY` and `OPENAI_BASE_URL` environment variables. For custom authentication, instantiate your own `openai.OpenAI` or `openai.AsyncOpenAI` client with custom `default_headers` or `http_client` parameters and pass it to the `client` argument of `ResponsesApiModelHandler`.