Implementing Interrupt Handling for Turn-Taking in Speech-to-Speech: A Complete Technical Guide

The Hugging Face Speech-to-Speech library implements barge-in functionality by checking three conditions in AudioHandler.on_speech_started—active response status, VAD interrupt flags, and runtime configuration—before canceling the assistant's current turn via RealtimeService.response.finish_response.

This guide explores how the huggingface/speech-to-speech repository handles real-time conversation flow, specifically implementing interrupt handling for turn-taking that allows users to cut off the assistant mid-utterance. Understanding this pipeline requires examining the interaction between Voice Activity Detection (VAD) events, mutable runtime configuration, and the WebSocket audio handler that orchestrates cancellation.

Architecture Overview

The turn-taking system operates through a tightly-coupled real-time pipeline consisting of three distinct layers. When a user speaks while the assistant is responding, the system must detect the intrusion, evaluate whether interruption is permitted, and gracefully terminate the ongoing response before processing new input.

The implementation spans three critical components:

Voice Activity Detection and Event Emission

The interruption process begins in the VAD layer when the detector identifies the start of a new speech segment. The system creates a SpeechStartedEvent that carries metadata essential for turn-taking decisions.

The SpeechStartedEvent Structure

In src/speech_to_speech/pipeline/events.py (lines 30-37), the event class includes an interrupt_response flag defaulting to True:

class SpeechStartedEvent(PipelineEvent):
    """Emitted when VAD detects speech start"""
    event_type: str = "speech_started"
    interrupt_response: bool = True  # Controls whether to cancel current response

    timestamp: float

This event propagates through the text_output_queue to downstream handlers. The VAD emits this event regardless of whether the assistant is currently speaking, leaving the interruption decision to the handler layer.

Runtime Configuration for Interruption Control

The system respects session-level configuration that can disable interruption behavior dynamically. This allows clients to implement "polite" modes where the assistant always completes its thought.

RuntimeConfig.interrupt_response_enabled

Located in src/speech_to_speech/api/openai_realtime/runtime_config.py (lines 58-76), this property extracts the interruption preference from the session configuration:

@property
def interrupt_response_enabled(self) -> bool:
    """Check if interruption is enabled in session config"""
    config = self.session
    if hasattr(config, 'audio'):
        audio = config.audio
        if hasattr(audio, 'input'):
            input_config = audio.input
            if hasattr(input_config, 'turn_detection'):
                turn_detection = input_config.turn_detection
                if hasattr(turn_detection, 'interrupt_response'):
                    return turn_detection.interrupt_response
    return True  # Default: interruption enabled

The implementation handles both Pydantic models and plain dictionaries, falling back to True (OpenAI's default behavior) when the field is unspecified.

AudioHandler: Executing the Interruption

The AudioHandler class in src/speech_to_speech/api/openai_realtime/handlers/audio.py serves as the orchestration point for turn-taking logic. The on_speech_started method (lines 96-108) evaluates whether to cancel the active response.

Interruption Decision Logic

When processing a SpeechStartedEvent, the handler verifies three conditions:

  1. Active Response Check (st.in_response) – Confirms the assistant is currently generating a response
  2. Event Flag Check (event.interrupt_response) – Verifies the VAD marked this speech start as interruptible
  3. Config Check (st.runtime_config.interrupt_response_enabled) – Ensures the session permits interruptions

If all conditions satisfy, the handler executes cancellation:

def on_speech_started(self, event: SpeechStartedEvent, st: ConnectionState) -> None:
    """Handle speech start event with potential interruption"""
    if (st.in_response and 
        event.interrupt_response and 
        st.runtime_config.interrupt_response_enabled):
        
        # Cancel active response

        self.response.finish_response(
            status="cancelled", 
            reason="turn_detected"
        )
        
    # Always prepare for new input

    self.response._start_item()

The finish_response call propagates through RealtimeService, which clears the current response generation, discarding pending audio chunks and TTS output. This guarantees the assistant cannot speak over the user's new utterance.

Configuring Interrupt Handling

Clients control barge-in behavior through the turn_detection.interrupt_response field in session creation requests or updates.

Session Configuration Example

To enable interruption (barge-in):

{
  "audio": {
    "input": {
      "turn_detection": {
        "type": "server_vad",
        "interrupt_response": true
      }
    }
  }
}

Setting "interrupt_response": false disables interruption, forcing the assistant to complete its current turn even if the user speaks.

Practical Implementation Examples

Creating a Session with Interruption Enabled

import speech_to_speech as sts

session_cfg = {
    "audio": {
        "input": {
            "turn_detection": {
                "type": "server_vad",
                "interrupt_response": True
            }
        }
    }
}

service = sts.api.openai_realtime.realtime_service.RealtimeService()
service.runtime_config.apply_session_update(
    sts.api.openai_realtime.runtime_config.RealtimeSessionCreateRequest(**session_cfg)
)

Dynamically Disabling Interruption Mid-Session


# Disable interruption during runtime

service.runtime_config.session.audio.input.turn_detection.interrupt_response = False

# Now speech starts won't cancel active responses

Client-Side Event Listening

from speech_to_speech.api.openai_realtime.client import RealtimeClient

client = RealtimeClient(url="ws://localhost:8000/realtime")

def on_speech_started(event):
    print(f"User speaking - interrupt allowed: {event.interrupt_response}")

client.on("speech_started", on_speech_started)
client.connect()

Key Source Files Reference

Understanding interrupt handling for turn-taking requires familiarity with these specific files:

Summary

The Hugging Face Speech-to-Speech library implements robust turn-taking through coordinated event handling:

  • VAD Detection emits SpeechStartedEvent with an interrupt_response flag indicating whether the new speech should pre-empt ongoing responses
  • Runtime Configuration provides interrupt_response_enabled to dynamically control barge-in behavior per the OpenAI Realtime API specification
  • AudioHandler executes the cancellation through RealtimeService.response.finish_response when conditions permit, ensuring the assistant stops before processing new user input
  • Configuration occurs via audio.input.turn_detection.interrupt_response in session creation or update requests

Frequently Asked Questions

How does the system prevent the assistant from talking over the user?

When AudioHandler.on_speech_started detects valid interruption conditions, it calls response.finish_response(status="cancelled", reason="turn_detected"). This propagates through RealtimeService, which immediately terminates the current response generator and discards pending TTS audio chunks, ensuring the assistant stops speaking before the user's new utterance is processed.

Can I disable interrupt handling for specific parts of a conversation?

Yes. Modify runtime_config.session.audio.input.turn_detection.interrupt_response to False at any point during the session. Subsequent SpeechStartedEvent instances will not trigger cancellation even if the assistant is responding, creating a "polite" mode where the assistant always completes its current turn.

What happens if the VAD detects speech but the configuration has interruption disabled?

The AudioHandler checks st.runtime_config.interrupt_response_enabled before executing cancellation. If this returns False, the handler skips the finish_response call and proceeds directly to _start_item() for the new input. The assistant continues its current response while the new user audio buffers for processing afterward.

Where is the default interruption behavior defined?

The default value True for interrupt_response appears in two locations: the SpeechStartedEvent class definition in src/speech_to_speech/pipeline/events.py (line 33), and the fallback return value in RuntimeConfig.interrupt_response_enabled at line 76 of src/speech_to_speech/api/openai_realtime/runtime_config.py. This mirrors the OpenAI Realtime API default behavior.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →