How the Speech-to-Speech Pipeline Handles Interruption When the User Speaks While AI Is Responding
When the user speaks while the AI is responding, the pipeline cancels the active LLM and TTS streams via the on_speech_started handler in the audio module, then immediately begins processing the new user input.
The huggingface/speech-to-speech repository implements a real-time conversational AI system where users can naturally interrupt the AI mid-response. This "barge-in" capability relies on continuous audio monitoring and a coordinated cancellation protocol across the STT, LLM, and TTS components.
The Interruption Detection Flow
The pipeline continuously analyzes incoming audio through a Voice Activity Detector (VAD). When the VAD identifies new user speech during an active AI generation cycle, it triggers a structured interruption sequence.
Voice Activity Detection and SpeechStartedEvent
When the VAD detects speech onset while the system is in a response state (st.in_response), it emits a SpeechStartedEvent. This event carries an interrupt_response flag that signals the system's intent to cancel the ongoing generation.
The Cancellation Logic in audio.py
The primary handler resides in src/speech_to_speech/api/openai_realtime/handlers/audio.py. The on_speech_started method implements a three-gate check before cancelling:
st.in_responseindicates the system is currently generatingevent.interrupt_responseis set toTruein the eventst.runtime_config.interrupt_response_enabledpermits interruptions in the current session
Only when all three conditions are met does the handler invoke response.finish_response with status "cancelled" and reason "turn_detected".
def on_speech_started(self, conn_id: str, event: SpeechStartedEvent) -> list[ServerEvent]:
# Cancel the active response if interruptions are enabled
if st.in_response and event.interrupt_response and st.runtime_config.interrupt_response_enabled:
events.extend(response.finish_response(conn_id, status="cancelled", reason="turn_detected"))
Following the cancellation, the handler immediately initializes a new input item for the incoming speech via _start_input_item, ensuring the new user utterance enters the STT → LLM → TTS pipeline without delay.
input_item_id = self._start_input_item(
conn_id,
preserve_active_response=preserve_active_response,
)
Runtime Configuration for Interruption Control
The interruption feature is configurable per session through src/speech_to_speech/api/openai_realtime/runtime_config.py. The interrupt_response_enabled property dynamically reads the turn_detection.interrupt_response setting from the session configuration, defaulting to True when unspecified.
@property
def interrupt_response_enabled(self) -> bool:
# Reads `turn_detection.interrupt_response` from the session config
if hasattr(td, "interrupt_response"):
val = td.interrupt_response
else:
val = td.get("interrupt_response", True)
return bool(val)
This configuration allows developers to disable barge-in for specific use cases where uninterrupted AI delivery is required.
Generator-Level Cancellation
Beyond the API event layer, both language model and text-to-speech generators actively monitor interruption signals to terminate processing loops gracefully.
LLM Cancellation
The base language model implementation in src/speech_to_speech/LLM/base_openai_compatible_language_model.py checks the interrupted flag during generation. When detected, the loop logs the cancellation and exits immediately.
logger.info("LLM generation cancelled (interruption)")
TTS Cancellation
Similarly, TTS handlers such as src/speech_to_speech/TTS/qwen3_tts_handler.py monitor for interruptions to stop audio synthesis. Upon detection, the handler logs the cancellation event.
logger.info("TTS generation cancelled (interruption)")
This dual-layer cancellation ensures that computational resources are freed instantly when the user begins speaking.
Practical Implementation Examples
Developers can programmatically test or configure interruption behavior using the pipeline's public interfaces.
Manually Triggering an Interruption
For testing purposes, you can simulate a user interruption by constructing a SpeechStartedEvent with interrupt_response=True:
from speech_to_speech.api.openai_realtime.handlers.audio import AudioHandler
from speech_to_speech.api.openai_realtime.events import SpeechStartedEvent
handler = AudioHandler(service=my_service)
event = SpeechStartedEvent(
interrupt_response=True, # enable interruption
turn_id="turn-123",
turn_revision=1,
audio_start_ms=0,
)
handler.on_speech_started(conn_id="conn-1", event=event)
# The active response (if any) will be cancelled and a new input item created.
Disabling Interruption in Session Configuration
To prevent interruptions for specific sessions, set interrupt_response to False in the turn detection configuration:
session_config = {
"turn_detection": {"type": "server_vad", "interrupt_response": False}
}
runtime_cfg = RuntimeConfig(session=session_config)
assert not runtime_cfg.interrupt_response_enabled
# In this mode, SpeechStartedEvents will *not* cancel the ongoing response.
Summary
- Three-condition gate: Interruptions only occur when the system is responding, the event permits interruption, and the runtime configuration enables it.
- Centralized handler: The
on_speech_startedmethod insrc/speech_to_speech/api/openai_realtime/handlers/audio.pycoordinates cancellation and new input creation. - Configurable behavior: The
interrupt_response_enabledproperty inruntime_config.pyallows per-session control over barge-in capabilities. - Generator awareness: Both LLM and TTS components actively check interruption flags to ensure immediate termination of ongoing generation.
- Seamless transition: After cancellation, the pipeline immediately routes the new speech through
_start_input_itemfor transcription and response generation.
Frequently Asked Questions
What triggers an interruption in the speech-to-speech pipeline?
An interruption triggers when the Voice Activity Detector (VAD) identifies new user speech during an active AI response, emitting a SpeechStartedEvent with interrupt_response=True. The on_speech_started handler in audio.py then validates runtime permissions before cancelling the current generation.
Can interruptions be disabled for specific sessions?
Yes. Set turn_detection.interrupt_response to False in the session configuration. The interrupt_response_enabled property in runtime_config.py reads this value, and when False, the on_speech_started handler will not cancel active responses even when new speech is detected.
How do the LLM and TTS generators know when to stop?
Both generators monitor an interrupted flag within their processing loops. The LLM implementation in base_openai_compatible_language_model.py and TTS handlers like qwen3_tts_handler.py check this flag periodically, logging cancellation messages and exiting their generation loops immediately upon detection.
What happens to the partially generated audio when an interruption occurs?
When an interruption is confirmed, the response.finish_response method emits a ResponseFinishedEvent with status "cancelled", which signals the TTS stream to truncate. Any partially synthesized audio is discarded, and the system immediately begins processing the new user input through _start_input_item.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →