# How Interruption Handling and Barge-In Work in the Hugging Face Speech-to-Speech Pipeline

> Understand interruption handling and barge-in in the Hugging Face speech-to-speech pipeline. See how user speech cancels assistant output for a seamless experience.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: internals
- Published: 2026-07-08

---

**The speech-to-speech service implements OpenAI Realtime API semantics for barge-in by detecting user speech during assistant output, cancelling the active response via a generation-aware CancelScope, and filtering stale audio through a discard guard before cleanly resetting the state.**

The huggingface/speech-to-speech repository provides a real-time speech-to-speech pipeline that follows OpenAI Realtime API conventions. Interruption handling and barge-in allow users to speak over the assistant, triggering immediate cancellation of the ongoing response and preventing stale audio from reaching the client. This mechanism relies on three tightly coupled components that check runtime flags, detect speech events, and manage generation-scoped cancellation.

## Core Components Controlling Barge-In

Three primary components work together to detect interruptions and manage the cancellation lifecycle.

### RuntimeConfig.interrupt_response_enabled

In [`src/speech_to_speech/api/openai_realtime/runtime_config.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/api/openai_realtime/runtime_config.py), the `RuntimeConfig` class exposes the `interrupt_response_enabled` property. This mirrors the OpenAI `turn_detection.interrupt_response` flag and defaults to `True`. The flag is read from `session.audio.input.turn_detection` and determines whether a speech-started event is permitted to interrupt an active assistant response.

```python
class RuntimeConfig(BaseModel):
    @property
    def interrupt_response_enabled(self) -> bool:
        """Read `turn_detection.interrupt_response` from the session config.
        Defaults to True (OpenAI default)."""
        td = self.session.audio.input.turn_detection
        if td is None:
            return True
        # Handles both Pydantic model and plain dict cases

        return getattr(td, "interrupt_response", td.get("interrupt_response", True))

```

### AudioHandler.on_speech_started

In [`src/speech_to_speech/api/openai_realtime/handlers/audio.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/api/openai_realtime/handlers/audio.py), the `AudioHandler.on_speech_started` method receives `SpeechStartedEvent` from the VAD. When the connection state `st.in_response` is true and the event’s `interrupt_response` attribute is true, it calls `response.finish_response(..., status="cancelled")` to initiate the discard process.

```python
def on_speech_started(self, conn_id: str, event: SpeechStartedEvent) -> list[ServerEvent]:
    st = self._state(conn_id)
    # Cancel the active response if interrupts are enabled

    if st.in_response and event.interrupt_response and st.runtime_config.interrupt_response_enabled:
        events.extend(response.finish_response(conn_id, status="cancelled", reason="turn_detected"))
    # … continue handling the new input …

    return events

```

### CancelScope and Generation Tracking

The pipeline runs each unit inside an `anyio.CancelScope` managed in [`src/speech_to_speech/pipeline/unit.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/pipeline/unit.py). When `finish_response` is invoked, it sets `self.cancel_scope.discarding = True` and increments the generation counter. Pending audio and text items carrying the old generation ID are dropped, while new items pass through. The flag clears automatically when the `__RESPONSE_DONE__` sentinel (the `AUDIO_RESPONSE_DONE` marker) is processed.

## Step-by-Step Barge-In Flow

When a user speaks during assistant output, the system executes the following sequence:

1. **Assistant is speaking** – The pipeline has emitted a `ResponseCreatedEvent` and `ResponseAudioDeltaEvent` messages. The `CancelScope` for that unit is not discarding.

2. **VAD detects user speech** – A `SpeechStartedEvent` is placed onto the `text_output_queue`.

3. **Handler evaluates cancellation** – `AudioHandler.on_speech_started` checks three conditions: `st.in_response` is true, `event.interrupt_response` is true, and `st.runtime_config.interrupt_response_enabled` is true. When all pass, it calls `response.finish_response(conn_id, status="cancelled", reason="turn_detected")`.

4. **Cancellation scope activates** – `Response.finish_response` sets `self.cancel_scope.discarding = True` and bumps `self.cancel_scope.generation`.

5. **Pending items are filtered** – In [`src/speech_to_speech/api/openai_realtime/websocket_router.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/api/openai_realtime/websocket_router.py), the `_drain_pending_response_events` loop processes queued items. The `_generation_is_discardable` helper checks if text or audio events belong to the old generation. If true, the events are silently dropped.

```python
def _drain_pending_response_events(...):
    ...
    elif isinstance(item, AssistantTextEvent):
        # Skip text that belongs to a generation that is currently being discarded

        if _generation_is_discardable(unit, item.cancel_generation):
            continue
        events = unit.service.dispatch_pipeline_event(session_id, item)
        ...

```

6. **Sentinel clears the state** – When the original response’s final audio chunk arrives, the router recognizes the `__RESPONSE_DONE__` sentinel. It sends `response.output_audio.done` and `response.done` events, then clears the discarding flag (`self.cancel_scope.discarding = False`).

7. **New turn begins** – The next `SpeechStartedEvent` creates a fresh generation. All new audio and text events emit normally, providing the client with a clean continuation.

## Validation Through Testing

The test suite in [`tests/openai_realtime/test_websocket_router.py`](https://github.com/huggingface/speech-to-speech/blob/main/tests/openai_realtime/test_websocket_router.py) verifies this behavior:

- `test_barge_in_discard_clears_after_response_done` confirms that after a barge-in, the `discarding` flag is set and later cleared once the response-done sentinel processes.
- `test_speech_started_cancels_pending_implicit_response` validates that implicit responses (generated by VAD → STT → LLM → TTS chains) cancel immediately when user speech is detected.

## Summary

- **Barge-in is opt-in via configuration** – The `interrupt_response_enabled` flag in `RuntimeConfig` controls whether user speech can interrupt assistant output, defaulting to `True` to match OpenAI behavior.
- **Cancellation is generation-scoped** – `anyio.CancelScope` tracks generations; setting `discarding = True` filters stale pipeline events while the old response winds down.
- **State clears automatically** – The `__RESPONSE_DONE__` sentinel triggers cleanup, ensuring the discarding flag does not persist across turns.
- **Stale output is blocked** – The `_generation_is_discardable` guard in the WebSocket router prevents old audio deltas from reaching the client after an interruption.

## Frequently Asked Questions

### What triggers a barge-in cancellation in the speech-to-speech pipeline?

A barge-in triggers when the VAD emits a `SpeechStartedEvent` while `st.in_response` is true, the event’s `interrupt_response` attribute is true, and `RuntimeConfig.interrupt_response_enabled` returns true. This combination causes `AudioHandler.on_speech_started` to invoke `finish_response` with status `"cancelled"`.

### How does the system prevent old audio from playing after an interruption?

The system uses `_generation_is_discardable` in [`src/speech_to_speech/api/openai_realtime/websocket_router.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/api/openai_realtime/websocket_router.py) to check if queued audio or text events belong to a generation marked for discarding. Events from the cancelled generation are dropped; only events from the new generation are emitted to the client.

### When does the discarding state reset after a barge-in?

The discarding state resets when the `__RESPONSE_DONE__` sentinel (or `AUDIO_RESPONSE_DONE` marker) is processed. This sentinel signals the end of the original response's audio stream, allowing the router to set `self.cancel_scope.discarding = False` and prepare for the next turn.

### Can developers disable barge-in functionality?

Yes. Developers can disable interruption handling by setting `turn_detection.interrupt_response` to `False` in the session configuration. This causes `RuntimeConfig.interrupt_response_enabled` to return `False`, preventing `AudioHandler.on_speech_started` from cancelling active responses when user speech is detected.