Implementing Interrupt Handling for Turn-Taking in Speech-to-Speech: A Complete Technical Guide
The Hugging Face Speech-to-Speech library implements barge-in functionality by checking three conditions in AudioHandler.on_speech_started—active response status, VAD interrupt flags, and runtime configuration—before canceling the assistant's current turn via RealtimeService.response.finish_response.
This guide explores how the huggingface/speech-to-speech repository handles real-time conversation flow, specifically implementing interrupt handling for turn-taking that allows users to cut off the assistant mid-utterance. Understanding this pipeline requires examining the interaction between Voice Activity Detection (VAD) events, mutable runtime configuration, and the WebSocket audio handler that orchestrates cancellation.
Architecture Overview
The turn-taking system operates through a tightly-coupled real-time pipeline consisting of three distinct layers. When a user speaks while the assistant is responding, the system must detect the intrusion, evaluate whether interruption is permitted, and gracefully terminate the ongoing response before processing new input.
The implementation spans three critical components:
- Pipeline Events (
src/speech_to_speech/pipeline/events.py) – Typed event carriers that travel through the processing queue - Runtime Configuration (
src/speech_to_speech/api/openai_realtime/runtime_config.py) – Mutable session settings controlling interruption behavior - Audio Handler (
src/speech_to_speech/api/openai_realtime/handlers/audio.py) – The decision engine that executes cancellation logic
Voice Activity Detection and Event Emission
The interruption process begins in the VAD layer when the detector identifies the start of a new speech segment. The system creates a SpeechStartedEvent that carries metadata essential for turn-taking decisions.
The SpeechStartedEvent Structure
In src/speech_to_speech/pipeline/events.py (lines 30-37), the event class includes an interrupt_response flag defaulting to True:
class SpeechStartedEvent(PipelineEvent):
"""Emitted when VAD detects speech start"""
event_type: str = "speech_started"
interrupt_response: bool = True # Controls whether to cancel current response
timestamp: float
This event propagates through the text_output_queue to downstream handlers. The VAD emits this event regardless of whether the assistant is currently speaking, leaving the interruption decision to the handler layer.
Runtime Configuration for Interruption Control
The system respects session-level configuration that can disable interruption behavior dynamically. This allows clients to implement "polite" modes where the assistant always completes its thought.
RuntimeConfig.interrupt_response_enabled
Located in src/speech_to_speech/api/openai_realtime/runtime_config.py (lines 58-76), this property extracts the interruption preference from the session configuration:
@property
def interrupt_response_enabled(self) -> bool:
"""Check if interruption is enabled in session config"""
config = self.session
if hasattr(config, 'audio'):
audio = config.audio
if hasattr(audio, 'input'):
input_config = audio.input
if hasattr(input_config, 'turn_detection'):
turn_detection = input_config.turn_detection
if hasattr(turn_detection, 'interrupt_response'):
return turn_detection.interrupt_response
return True # Default: interruption enabled
The implementation handles both Pydantic models and plain dictionaries, falling back to True (OpenAI's default behavior) when the field is unspecified.
AudioHandler: Executing the Interruption
The AudioHandler class in src/speech_to_speech/api/openai_realtime/handlers/audio.py serves as the orchestration point for turn-taking logic. The on_speech_started method (lines 96-108) evaluates whether to cancel the active response.
Interruption Decision Logic
When processing a SpeechStartedEvent, the handler verifies three conditions:
- Active Response Check (
st.in_response) – Confirms the assistant is currently generating a response - Event Flag Check (
event.interrupt_response) – Verifies the VAD marked this speech start as interruptible - Config Check (
st.runtime_config.interrupt_response_enabled) – Ensures the session permits interruptions
If all conditions satisfy, the handler executes cancellation:
def on_speech_started(self, event: SpeechStartedEvent, st: ConnectionState) -> None:
"""Handle speech start event with potential interruption"""
if (st.in_response and
event.interrupt_response and
st.runtime_config.interrupt_response_enabled):
# Cancel active response
self.response.finish_response(
status="cancelled",
reason="turn_detected"
)
# Always prepare for new input
self.response._start_item()
The finish_response call propagates through RealtimeService, which clears the current response generation, discarding pending audio chunks and TTS output. This guarantees the assistant cannot speak over the user's new utterance.
Configuring Interrupt Handling
Clients control barge-in behavior through the turn_detection.interrupt_response field in session creation requests or updates.
Session Configuration Example
To enable interruption (barge-in):
{
"audio": {
"input": {
"turn_detection": {
"type": "server_vad",
"interrupt_response": true
}
}
}
}
Setting "interrupt_response": false disables interruption, forcing the assistant to complete its current turn even if the user speaks.
Practical Implementation Examples
Creating a Session with Interruption Enabled
import speech_to_speech as sts
session_cfg = {
"audio": {
"input": {
"turn_detection": {
"type": "server_vad",
"interrupt_response": True
}
}
}
}
service = sts.api.openai_realtime.realtime_service.RealtimeService()
service.runtime_config.apply_session_update(
sts.api.openai_realtime.runtime_config.RealtimeSessionCreateRequest(**session_cfg)
)
Dynamically Disabling Interruption Mid-Session
# Disable interruption during runtime
service.runtime_config.session.audio.input.turn_detection.interrupt_response = False
# Now speech starts won't cancel active responses
Client-Side Event Listening
from speech_to_speech.api.openai_realtime.client import RealtimeClient
client = RealtimeClient(url="ws://localhost:8000/realtime")
def on_speech_started(event):
print(f"User speaking - interrupt allowed: {event.interrupt_response}")
client.on("speech_started", on_speech_started)
client.connect()
Key Source Files Reference
Understanding interrupt handling for turn-taking requires familiarity with these specific files:
src/speech_to_speech/pipeline/events.py– DefinesSpeechStartedEventwith theinterrupt_responseflagsrc/speech_to_speech/api/openai_realtime/runtime_config.py– ImplementsRuntimeConfig.interrupt_response_enabledpropertysrc/speech_to_speech/api/openai_realtime/handlers/audio.py– ContainsAudioHandler.on_speech_startedcancellation logictests/test_speculative_turns.py– Unit tests verifying VAD-driven interruptions respect configuration flagsscripts/listen_and_play_realtime.py– Example client implementation demonstrating session configuration
Summary
The Hugging Face Speech-to-Speech library implements robust turn-taking through coordinated event handling:
- VAD Detection emits
SpeechStartedEventwith aninterrupt_responseflag indicating whether the new speech should pre-empt ongoing responses - Runtime Configuration provides
interrupt_response_enabledto dynamically control barge-in behavior per the OpenAI Realtime API specification - AudioHandler executes the cancellation through
RealtimeService.response.finish_responsewhen conditions permit, ensuring the assistant stops before processing new user input - Configuration occurs via
audio.input.turn_detection.interrupt_responsein session creation or update requests
Frequently Asked Questions
How does the system prevent the assistant from talking over the user?
When AudioHandler.on_speech_started detects valid interruption conditions, it calls response.finish_response(status="cancelled", reason="turn_detected"). This propagates through RealtimeService, which immediately terminates the current response generator and discards pending TTS audio chunks, ensuring the assistant stops speaking before the user's new utterance is processed.
Can I disable interrupt handling for specific parts of a conversation?
Yes. Modify runtime_config.session.audio.input.turn_detection.interrupt_response to False at any point during the session. Subsequent SpeechStartedEvent instances will not trigger cancellation even if the assistant is responding, creating a "polite" mode where the assistant always completes its current turn.
What happens if the VAD detects speech but the configuration has interruption disabled?
The AudioHandler checks st.runtime_config.interrupt_response_enabled before executing cancellation. If this returns False, the handler skips the finish_response call and proceeds directly to _start_item() for the new input. The assistant continues its current response while the new user audio buffers for processing afterward.
Where is the default interruption behavior defined?
The default value True for interrupt_response appears in two locations: the SpeechStartedEvent class definition in src/speech_to_speech/pipeline/events.py (line 33), and the fallback return value in RuntimeConfig.interrupt_response_enabled at line 76 of src/speech_to_speech/api/openai_realtime/runtime_config.py. This mirrors the OpenAI Realtime API default behavior.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →