How Interruption Handling and Barge-In Work in the Hugging Face Speech-to-Speech Pipeline
The speech-to-speech service implements OpenAI Realtime API semantics for barge-in by detecting user speech during assistant output, cancelling the active response via a generation-aware CancelScope, and filtering stale audio through a discard guard before cleanly resetting the state.
The huggingface/speech-to-speech repository provides a real-time speech-to-speech pipeline that follows OpenAI Realtime API conventions. Interruption handling and barge-in allow users to speak over the assistant, triggering immediate cancellation of the ongoing response and preventing stale audio from reaching the client. This mechanism relies on three tightly coupled components that check runtime flags, detect speech events, and manage generation-scoped cancellation.
Core Components Controlling Barge-In
Three primary components work together to detect interruptions and manage the cancellation lifecycle.
RuntimeConfig.interrupt_response_enabled
In src/speech_to_speech/api/openai_realtime/runtime_config.py, the RuntimeConfig class exposes the interrupt_response_enabled property. This mirrors the OpenAI turn_detection.interrupt_response flag and defaults to True. The flag is read from session.audio.input.turn_detection and determines whether a speech-started event is permitted to interrupt an active assistant response.
class RuntimeConfig(BaseModel):
@property
def interrupt_response_enabled(self) -> bool:
"""Read `turn_detection.interrupt_response` from the session config.
Defaults to True (OpenAI default)."""
td = self.session.audio.input.turn_detection
if td is None:
return True
# Handles both Pydantic model and plain dict cases
return getattr(td, "interrupt_response", td.get("interrupt_response", True))
AudioHandler.on_speech_started
In src/speech_to_speech/api/openai_realtime/handlers/audio.py, the AudioHandler.on_speech_started method receives SpeechStartedEvent from the VAD. When the connection state st.in_response is true and the event’s interrupt_response attribute is true, it calls response.finish_response(..., status="cancelled") to initiate the discard process.
def on_speech_started(self, conn_id: str, event: SpeechStartedEvent) -> list[ServerEvent]:
st = self._state(conn_id)
# Cancel the active response if interrupts are enabled
if st.in_response and event.interrupt_response and st.runtime_config.interrupt_response_enabled:
events.extend(response.finish_response(conn_id, status="cancelled", reason="turn_detected"))
# … continue handling the new input …
return events
CancelScope and Generation Tracking
The pipeline runs each unit inside an anyio.CancelScope managed in src/speech_to_speech/pipeline/unit.py. When finish_response is invoked, it sets self.cancel_scope.discarding = True and increments the generation counter. Pending audio and text items carrying the old generation ID are dropped, while new items pass through. The flag clears automatically when the __RESPONSE_DONE__ sentinel (the AUDIO_RESPONSE_DONE marker) is processed.
Step-by-Step Barge-In Flow
When a user speaks during assistant output, the system executes the following sequence:
-
Assistant is speaking – The pipeline has emitted a
ResponseCreatedEventandResponseAudioDeltaEventmessages. TheCancelScopefor that unit is not discarding. -
VAD detects user speech – A
SpeechStartedEventis placed onto thetext_output_queue. -
Handler evaluates cancellation –
AudioHandler.on_speech_startedchecks three conditions:st.in_responseis true,event.interrupt_responseis true, andst.runtime_config.interrupt_response_enabledis true. When all pass, it callsresponse.finish_response(conn_id, status="cancelled", reason="turn_detected"). -
Cancellation scope activates –
Response.finish_responsesetsself.cancel_scope.discarding = Trueand bumpsself.cancel_scope.generation. -
Pending items are filtered – In
src/speech_to_speech/api/openai_realtime/websocket_router.py, the_drain_pending_response_eventsloop processes queued items. The_generation_is_discardablehelper checks if text or audio events belong to the old generation. If true, the events are silently dropped.
def _drain_pending_response_events(...):
...
elif isinstance(item, AssistantTextEvent):
# Skip text that belongs to a generation that is currently being discarded
if _generation_is_discardable(unit, item.cancel_generation):
continue
events = unit.service.dispatch_pipeline_event(session_id, item)
...
-
Sentinel clears the state – When the original response’s final audio chunk arrives, the router recognizes the
__RESPONSE_DONE__sentinel. It sendsresponse.output_audio.doneandresponse.doneevents, then clears the discarding flag (self.cancel_scope.discarding = False). -
New turn begins – The next
SpeechStartedEventcreates a fresh generation. All new audio and text events emit normally, providing the client with a clean continuation.
Validation Through Testing
The test suite in tests/openai_realtime/test_websocket_router.py verifies this behavior:
test_barge_in_discard_clears_after_response_doneconfirms that after a barge-in, thediscardingflag is set and later cleared once the response-done sentinel processes.test_speech_started_cancels_pending_implicit_responsevalidates that implicit responses (generated by VAD → STT → LLM → TTS chains) cancel immediately when user speech is detected.
Summary
- Barge-in is opt-in via configuration – The
interrupt_response_enabledflag inRuntimeConfigcontrols whether user speech can interrupt assistant output, defaulting toTrueto match OpenAI behavior. - Cancellation is generation-scoped –
anyio.CancelScopetracks generations; settingdiscarding = Truefilters stale pipeline events while the old response winds down. - State clears automatically – The
__RESPONSE_DONE__sentinel triggers cleanup, ensuring the discarding flag does not persist across turns. - Stale output is blocked – The
_generation_is_discardableguard in the WebSocket router prevents old audio deltas from reaching the client after an interruption.
Frequently Asked Questions
What triggers a barge-in cancellation in the speech-to-speech pipeline?
A barge-in triggers when the VAD emits a SpeechStartedEvent while st.in_response is true, the event’s interrupt_response attribute is true, and RuntimeConfig.interrupt_response_enabled returns true. This combination causes AudioHandler.on_speech_started to invoke finish_response with status "cancelled".
How does the system prevent old audio from playing after an interruption?
The system uses _generation_is_discardable in src/speech_to_speech/api/openai_realtime/websocket_router.py to check if queued audio or text events belong to a generation marked for discarding. Events from the cancelled generation are dropped; only events from the new generation are emitted to the client.
When does the discarding state reset after a barge-in?
The discarding state resets when the __RESPONSE_DONE__ sentinel (or AUDIO_RESPONSE_DONE marker) is processed. This sentinel signals the end of the original response's audio stream, allowing the router to set self.cancel_scope.discarding = False and prepare for the next turn.
Can developers disable barge-in functionality?
Yes. Developers can disable interruption handling by setting turn_detection.interrupt_response to False in the session configuration. This causes RuntimeConfig.interrupt_response_enabled to return False, preventing AudioHandler.on_speech_started from cancelling active responses when user speech is detected.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →