How to Implement Response Streaming with Responses API Backend in Speech-to-Speech
Enable real-time token delivery by setting responses_api_stream=True in ResponsesApiLanguageModelHandlerArguments, which propagates the stream flag through the OpenAI-compatible client and yields incremental TextDelta events as the provider generates them.
The huggingface/speech-to-speech repository provides a production-ready pipeline for real-time voice conversations, supporting OpenAI-compatible Responses API backends from providers like Together, OpenAI, and Azure. Understanding how to implement response streaming with the Responses API backend is essential for building low-latency applications that display text as it generates rather than waiting for complete responses. This guide examines the three-layer architecture that controls streaming behavior and provides executable code samples for both high-level pipeline integration and low-level handler manipulation.
Understanding the Streaming Architecture
Response streaming operates through three coordinated layers in the codebase, each handling a specific aspect of the data flow from configuration to event consumption.
Configuration Layer
The responses_api_stream parameter in ResponsesApiLanguageModelHandlerArguments serves as the master switch for streaming behavior. Located in src/speech_to_speech/arguments_classes/responses_api_language_model_arguments.py at lines 21-27, this dataclass field defaults to True:
# src/speech_to_speech/arguments_classes/responses_api_language_model_arguments.py
responses_api_stream: bool = field(
default=True,
metadata={"help": "Enable continuous token‑by‑token delivery via the Responses API."},
)
When you instantiate the pipeline with --language-model-type responses-api, the s2s_pipeline.py file passes this value to the handler constructor at lines 855-861, ensuring the boolean propagates through the entire request lifecycle.
Request Layer
The ResponsesApiModelHandler._request method forwards the stream flag directly to the underlying OpenAI client. In src/speech_to_speech/LLM/responses_api_language_model.py at lines 97-104, the method signature binds the instance attribute self.stream to the API call:
# src/speech_to_speech/LLM/responses_api_language_model.py
def _request(self, api_input, optional_kwargs):
return self.client.responses.create(
model=self.model_name,
input=api_input,
stream=self.stream,
extra_body=self._extra_body,
timeout=self.request_timeout,
**optional_kwargs,
)
Setting stream=True causes the client to return an openai.Stream object rather than a blocking response object, establishing the HTTP connection for server-sent events.
Event Iteration Layer
When streaming is enabled, the handler processes the raw stream through _iter_stream_events, converting low-level SDK events into high-level ProviderEvent objects. This method in src/speech_to_speech/LLM/responses_api_language_model.py (lines 13-30) yields TextDelta instances containing incremental text fragments:
# src/speech_to_speech/LLM/responses_api_language_model.py
def _iter_stream_events(self, api_response: Stream) -> Iterator[ProviderEvent]:
for raw_event in api_response:
if isinstance(raw_event, ResponseTextDeltaEvent):
yield TextDelta(text=raw_event.delta)
elif isinstance(raw_event, ResponseOutputItemDoneEvent):
...
The pipeline consumes these events asynchronously, enabling real-time display of generated content without buffering the entire response.
Configuring Response Streaming
You can enable streaming through command-line arguments or programmatic configuration depending on your deployment requirements.
Command-line activation:
python -m speech_to_speech \
--language-model-type responses-api \
--responses_api_stream True \
--responses_api_language_model_name "gpt-5.4-mini"
Programmatic configuration:
from speech_to_speech.s2s_pipeline import SpeechToSpeechPipeline
from speech_to_speech.arguments_classes.responses_api_language_model_arguments import (
ResponsesApiLanguageModelHandlerArguments,
)
args = ResponsesApiLanguageModelHandlerArguments(
model_name="gpt-5.4-mini",
responses_api_stream=True,
)
pipeline = SpeechToSpeechPipeline(
language_model_type="responses-api",
responses_api_language_model_handler_kwargs=args,
)
Practical Implementation Examples
High-Level Pipeline Usage
This minimal script demonstrates streaming responses using the high-level pipeline interface:
# demo_stream_responses.py
import argparse
from speech_to_speech.s2s_pipeline import SpeechToSpeechPipeline
from speech_to_speech.arguments_classes.responses_api_language_model_arguments import (
ResponsesApiLanguageModelHandlerArguments,
)
parser = argparse.ArgumentParser()
parser.add_argument("--prompt", default="Tell me a short joke.")
args = parser.parse_args()
pipeline = SpeechToSpeechPipeline(
language_model_type="responses-api",
responses_api_language_model_handler_kwargs=ResponsesApiLanguageModelHandlerArguments(
model_name="gpt-5.4-mini",
responses_api_stream=True,
responses_api_disable_thinking=False,
),
)
for event in pipeline.run(prompt=args.prompt):
if isinstance(event, pipeline.provider_events.TextDelta):
print(event.text, end="", flush=True)
Running this script prints tokens as they arrive from the provider, achieving true real-time streaming behavior suitable for live chat interfaces.
Direct Handler Control
For fine-grained control over the streaming process, interact directly with ResponsesApiModelHandler:
# direct_handler_stream.py
from speech_to_speech.LLM.responses_api_language_model import ResponsesApiModelHandler
from speech_to_speech.arguments_classes.responses_api_language_model_arguments import (
ResponsesApiLanguageModelHandlerArguments,
)
args = ResponsesApiLanguageModelHandlerArguments(
model_name="gpt-5.4-mini",
responses_api_stream=True,
responses_api_disable_thinking=True,
)
handler = ResponsesApiModelHandler(
model_name=args.model_name,
stream=args.responses_api_stream,
disable_thinking=args.responses_api_disable_thinking,
client=None, # Uses OPENAI_API_KEY and BASE_URL env vars
)
api_input = [
{"type": "message", "role": "system", "content": [{"type": "input_text", "text": "You are a helpful assistant"}]},
{"type": "message", "role": "user", "content": [{"type": "input_text", "text": "What's the weather like today?"}]},
]
stream = handler._request(api_input, optional_kwargs={})
for ev in handler._iter_stream_events(stream):
if isinstance(ev, handler.TextDelta):
print(ev.text, end="", flush=True)
elif isinstance(ev, handler.ToolCall):
pass # Handle function calls dynamically
This approach exposes the raw openai.Stream and allows custom processing of ResponseTextDeltaEvent objects before they reach the pipeline.
Disabling Provider-Side Reasoning
Some providers emit "thinking" tokens alongside content. Disable this to stream only final output tokens:
args = ResponsesApiLanguageModelHandlerArguments(
model_name="gpt-5.4-mini",
responses_api_stream=True,
responses_api_disable_thinking=True, # Disables chain-of-thought
)
When responses_api_disable_thinking is set to True, the handler injects chat_template_kwargs={"enable_thinking": False} into the request payload via the logic in src/speech_to_speech/LLM/responses_api_language_model.py lines 86-94.
Summary
- Three-layer control: Streaming configuration flows from
responses_api_streamin the arguments dataclass, through the_requestmethod'sstreamparameter, to the_iter_stream_eventsgenerator that yieldsTextDeltaobjects. - Real-time delivery: The
ResponsesApiModelHandlerconverts OpenAIStreamobjects into consumable events, printing tokens immediately as the LLM generates them. - Thinking suppression: Use
responses_api_disable_thinkingto eliminate reasoning tokens and reduce bandwidth when only final output is required. - Flexible integration: The pipeline supports both high-level
SpeechToSpeechPipelineusage for standard workflows and directResponsesApiModelHandlermanipulation for custom event processing.
Frequently Asked Questions
Which providers support the Responses API streaming mode?
The implementation works with any OpenAI-compatible endpoint, including Together AI, OpenAI, and Azure OpenAI Service. As long as the provider supports the Responses API schema and streaming parameters, the ResponsesApiModelHandler will correctly yield TextDelta events according to the huggingface/speech-to-speech source code.
How do I disable streaming to receive complete responses?
Set responses_api_stream=False in your ResponsesApiLanguageModelHandlerArguments configuration. When disabled, the _request method passes stream=False to the client, causing responses.create() to return a complete response object rather than a Stream, and the handler skips the _iter_stream_events generator in favor of synchronous result processing.
What is the difference between TextDelta and ResponseOutputItemDoneEvent?
TextDelta represents incremental content fragments delivered via ResponseTextDeltaEvent during active streaming, suitable for displaying partial text immediately. ResponseOutputItemDoneEvent signals the completion of an output item, indicating that the provider has finished generating that specific content block and no further deltas will arrive for that item.
Can I use custom authentication headers with the streaming handler?
Yes. When initializing ResponsesApiModelHandler with client=None, the handler uses the default OpenAI client initialization which respects OPENAI_API_KEY and OPENAI_BASE_URL environment variables. For custom authentication, instantiate your own openai.OpenAI or openai.AsyncOpenAI client with custom default_headers or http_client parameters and pass it to the client argument of ResponsesApiModelHandler.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →