How Interruption Handling and Barge-In Work in the Hugging Face Speech-to-Speech Pipeline

The speech-to-speech service implements OpenAI Realtime API semantics for barge-in by detecting user speech during assistant output, cancelling the active response via a generation-aware CancelScope, and filtering stale audio through a discard guard before cleanly resetting the state.

The huggingface/speech-to-speech repository provides a real-time speech-to-speech pipeline that follows OpenAI Realtime API conventions. Interruption handling and barge-in allow users to speak over the assistant, triggering immediate cancellation of the ongoing response and preventing stale audio from reaching the client. This mechanism relies on three tightly coupled components that check runtime flags, detect speech events, and manage generation-scoped cancellation.

Core Components Controlling Barge-In

Three primary components work together to detect interruptions and manage the cancellation lifecycle.

RuntimeConfig.interrupt_response_enabled

In src/speech_to_speech/api/openai_realtime/runtime_config.py, the RuntimeConfig class exposes the interrupt_response_enabled property. This mirrors the OpenAI turn_detection.interrupt_response flag and defaults to True. The flag is read from session.audio.input.turn_detection and determines whether a speech-started event is permitted to interrupt an active assistant response.

class RuntimeConfig(BaseModel):
    @property
    def interrupt_response_enabled(self) -> bool:
        """Read `turn_detection.interrupt_response` from the session config.
        Defaults to True (OpenAI default)."""
        td = self.session.audio.input.turn_detection
        if td is None:
            return True
        # Handles both Pydantic model and plain dict cases

        return getattr(td, "interrupt_response", td.get("interrupt_response", True))

AudioHandler.on_speech_started

In src/speech_to_speech/api/openai_realtime/handlers/audio.py, the AudioHandler.on_speech_started method receives SpeechStartedEvent from the VAD. When the connection state st.in_response is true and the event’s interrupt_response attribute is true, it calls response.finish_response(..., status="cancelled") to initiate the discard process.

def on_speech_started(self, conn_id: str, event: SpeechStartedEvent) -> list[ServerEvent]:
    st = self._state(conn_id)
    # Cancel the active response if interrupts are enabled

    if st.in_response and event.interrupt_response and st.runtime_config.interrupt_response_enabled:
        events.extend(response.finish_response(conn_id, status="cancelled", reason="turn_detected"))
    # … continue handling the new input …

    return events

CancelScope and Generation Tracking

The pipeline runs each unit inside an anyio.CancelScope managed in src/speech_to_speech/pipeline/unit.py. When finish_response is invoked, it sets self.cancel_scope.discarding = True and increments the generation counter. Pending audio and text items carrying the old generation ID are dropped, while new items pass through. The flag clears automatically when the __RESPONSE_DONE__ sentinel (the AUDIO_RESPONSE_DONE marker) is processed.

Step-by-Step Barge-In Flow

When a user speaks during assistant output, the system executes the following sequence:

  1. Assistant is speaking – The pipeline has emitted a ResponseCreatedEvent and ResponseAudioDeltaEvent messages. The CancelScope for that unit is not discarding.

  2. VAD detects user speech – A SpeechStartedEvent is placed onto the text_output_queue.

  3. Handler evaluates cancellation – AudioHandler.on_speech_started checks three conditions: st.in_response is true, event.interrupt_response is true, and st.runtime_config.interrupt_response_enabled is true. When all pass, it calls response.finish_response(conn_id, status="cancelled", reason="turn_detected").

  4. Cancellation scope activates – Response.finish_response sets self.cancel_scope.discarding = True and bumps self.cancel_scope.generation.

  5. Pending items are filtered – In src/speech_to_speech/api/openai_realtime/websocket_router.py, the _drain_pending_response_events loop processes queued items. The _generation_is_discardable helper checks if text or audio events belong to the old generation. If true, the events are silently dropped.

def _drain_pending_response_events(...):
    ...
    elif isinstance(item, AssistantTextEvent):
        # Skip text that belongs to a generation that is currently being discarded

        if _generation_is_discardable(unit, item.cancel_generation):
            continue
        events = unit.service.dispatch_pipeline_event(session_id, item)
        ...
  1. Sentinel clears the state – When the original response’s final audio chunk arrives, the router recognizes the __RESPONSE_DONE__ sentinel. It sends response.output_audio.done and response.done events, then clears the discarding flag (self.cancel_scope.discarding = False).

  2. New turn begins – The next SpeechStartedEvent creates a fresh generation. All new audio and text events emit normally, providing the client with a clean continuation.

Validation Through Testing

The test suite in tests/openai_realtime/test_websocket_router.py verifies this behavior:

  • test_barge_in_discard_clears_after_response_done confirms that after a barge-in, the discarding flag is set and later cleared once the response-done sentinel processes.
  • test_speech_started_cancels_pending_implicit_response validates that implicit responses (generated by VAD → STT → LLM → TTS chains) cancel immediately when user speech is detected.

Summary

  • Barge-in is opt-in via configuration – The interrupt_response_enabled flag in RuntimeConfig controls whether user speech can interrupt assistant output, defaulting to True to match OpenAI behavior.
  • Cancellation is generation-scoped – anyio.CancelScope tracks generations; setting discarding = True filters stale pipeline events while the old response winds down.
  • State clears automatically – The __RESPONSE_DONE__ sentinel triggers cleanup, ensuring the discarding flag does not persist across turns.
  • Stale output is blocked – The _generation_is_discardable guard in the WebSocket router prevents old audio deltas from reaching the client after an interruption.

Frequently Asked Questions

What triggers a barge-in cancellation in the speech-to-speech pipeline?

A barge-in triggers when the VAD emits a SpeechStartedEvent while st.in_response is true, the event’s interrupt_response attribute is true, and RuntimeConfig.interrupt_response_enabled returns true. This combination causes AudioHandler.on_speech_started to invoke finish_response with status "cancelled".

How does the system prevent old audio from playing after an interruption?

The system uses _generation_is_discardable in src/speech_to_speech/api/openai_realtime/websocket_router.py to check if queued audio or text events belong to a generation marked for discarding. Events from the cancelled generation are dropped; only events from the new generation are emitted to the client.

When does the discarding state reset after a barge-in?

The discarding state resets when the __RESPONSE_DONE__ sentinel (or AUDIO_RESPONSE_DONE marker) is processed. This sentinel signals the end of the original response's audio stream, allowing the router to set self.cancel_scope.discarding = False and prepare for the next turn.

Can developers disable barge-in functionality?

Yes. Developers can disable interruption handling by setting turn_detection.interrupt_response to False in the session configuration. This causes RuntimeConfig.interrupt_response_enabled to return False, preventing AudioHandler.on_speech_started from cancelling active responses when user speech is detected.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →