# How response.cancel Interrupts Generation in the Speech-to-Speech Realtime Engine

> Understand how response.cancel interrupts generation in the huggingface speech-to-speech engine. Learn how it aborts LLM and TTS generation for instant system reset.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: internals
- Published: 2026-07-31

---

**The `response.cancel` mechanism aborts active LLM and TTS generation by incrementing a generation counter and setting a discarding flag, which pipeline threads consult to drop stale output and immediately restore the system to a listening state.**

The `huggingface/speech-to-speech` repository implements a realtime audio pipeline that requires robust interruption capabilities. When users need to stop an ongoing response, the `response.cancel` event provides immediate termination of generation and flushes pending audio queues through a coordinated cancellation protocol.

## The Cancellation Architecture

The interruption flow centers on three core components that work together to safely abort generation without corrupting pipeline state.

### CancelScope State Management

The `CancelScope` class in [`src/speech_to_speech/pipeline/cancel_scope.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/pipeline/cancel_scope.py) maintains a **generation counter** (`_gen`) and a **discarding flag** (`_discarding`). These properties allow pipeline threads to determine whether their output is still valid or belongs to a cancelled generation.

### WebSocket Router Handling

The `_dispatch_client_event` function in [`src/speech_to_speech/api/openai_realtime/websocket_router.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/api/openai_realtime/websocket_router.py) receives the `ResponseCancelEvent` and triggers the cancellation sequence. It invokes `unit.cancel_scope.cancel()`, flushes pending queues, and forwards the event to the service layer.

### ResponseHandler Finalization

The `handle_response_cancel` method in [`src/speech_to_speech/api/openai_realtime/handlers/response.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/api/openai_realtime/handlers/response.py) emits the `response.done` event with `status="cancelled"` and re-enables the `should_listen` event to restore voice activity detection.

## Step-by-Step Cancellation Flow

### 1. Client Sends the Cancel Event

The client transmits a JSON message over the WebSocket:

```json
{
  "type": "response.cancel"
}

```

### 2. Router Triggers CancelScope

In [`websocket_router.py`](https://github.com/huggingface/speech-to-speech/blob/main/websocket_router.py), the `_dispatch_client_event` function detects the `ResponseCancelEvent` and executes the following logic:

```python
elif isinstance(event, ResponseCancelEvent):
    was_active = service._state(session_id).in_response
    if was_active:
        unit.cancel_scope.cancel()          # ← bump generation, set discarding

    _flush_queue(unit.output_queue, preserve=_keep_audio_sentinel)
    _flush_queue(unit.text_output_queue, preserve=_keep_user_text_event)
    transport.discard_pending_audio()
    events = service.handle_response_cancel(session_id)
    await transport.send_events(events)
    unit.response_playing.clear()

```

This sequence performs four critical actions:
- **`unit.cancel_scope.cancel()`** increments the generation counter and sets `_discarding = True`
- **Queue flushing** removes buffered audio and text belonging to the cancelled response
- **`transport.discard_pending_audio()`** drops audio already sent to the client but not yet played
- **Event forwarding** notifies the service layer to finalize the response

### 3. Pipeline Threads Check Generation Counters

The `CancelScope` class implements the core cancellation semantics:

```python
class CancelScope:
    def __init__(self):
        self._gen = 0
        self._discarding = False
        self._discarded_generation = None

    def cancel(self) -> None:
        self._discarded_generation = self._gen
        self._gen = (self._gen + 1) & 0xFFFFFFFF   # wrap-around safety

        self._discarding = True

    def is_stale(self, gen: int) -> bool:
        return gen != self._gen

    @property
    def discarding(self) -> bool:
        return self._discarding

```

Each pipeline thread captures `cancel_scope.generation` at the start of processing. If `cancel()` has incremented the counter, `is_stale()` returns `True`, causing the thread to stop processing and drop its output.

### 4. Service Emits response.done

The `ResponseHandler.handle_response_cancel` method finalizes the cancelled response:

```python
def handle_response_cancel(self, conn_id: str) -> list[ServerEvent]:
    events = self.finish_response(conn_id, status="cancelled", reason="client_cancelled")
    should_listen = self._should_listen(conn_id)
    if should_listen:
        should_listen.set()                     # re-enable VAD listening

    logger.info("Response cancelled, listening re-enabled")
    return events

```

This emits `response.done` with `status="cancelled"` and sets the `should_listen` event to unblock the VAD.

## Implementation Details

### Send-Loop Guard Mechanism

Inside the per-unit send loop (`websocket_router._send_loop_for`), the code guards against emitting stale audio:

```python
if unit.cancel_scope.discarding and generation != unit.cancel_scope.generation:
    # Discard stale audio while a cancel is in effect

    continue

```

Once `cancel()` has been called, any audio generated with the old generation is ignored until a new `response.create` arrives.

### Reset on New Response

When a fresh `response.create` event arrives, the router invokes:

```python
unit.cancel_scope.new_response()

```

This clears the `discarding` flag and the discarded generation marker, allowing the next generation to flow normally.

## Practical Code Examples

### Server-Side Manual Cancellation

You can trigger cancellation programmatically from custom handlers:

```python
from speech_to_speech.pipeline.cancel_scope import CancelScope

def my_handler(unit):
    # Detect user interruption (e.g., hot-key)

    if user_pressed_stop():
        unit.cancel_scope.cancel()          # abort current generation

        # Push sentinel to unblock the send loop

        unit.output_queue.put(AudioOutput(audio=AUDIO_RESPONSE_DONE,
                                          cancel_generation=unit.cancel_scope.generation))

```

### Client-Side JavaScript Implementation

```javascript
// ws is an open WebSocket to /v1/realtime
function cancelCurrentResponse() {
  ws.send(JSON.stringify({type: "response.cancel"}));
}

```

### Testing the Cancel Flow

```python
def test_response_cancel_flushes_queues(setup):
    app, service, _, output_q, text_q, _, _, _, cancel_scope = setup
    ws.send_json({"type": "response.cancel"})
    assert cancel_scope.discarding          # flag is set

    # After the server sends response.done, the flag is cleared

    await wait_for(lambda: not cancel_scope.discarding)

```

## Key Source Files

| File | Purpose |
|------|---------|
| [`src/speech_to_speech/pipeline/cancel_scope.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/pipeline/cancel_scope.py) | Core cancellation state (generation counter, discarding flag) |
| [`src/speech_to_speech/api/openai_realtime/websocket_router.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/api/openai_realtime/websocket_router.py) | Entry point for `response.cancel`; triggers queue flushing |
| [`src/speech_to_speech/api/openai_realtime/handlers/response.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/api/openai_realtime/handlers/response.py) | Generates `response.done` with `status="cancelled"`; re-enables listening |
| [`tests/openai_realtime/test_websocket_router.py`](https://github.com/huggingface/speech-to-speech/blob/main/tests/openai_realtime/test_websocket_router.py) | Validates the cancel flow (discarding flag, queue flushing) |

## Summary

- **`response.cancel`** immediately interrupts LLM and TTS generation by incrementing a generation counter in `CancelScope`.
- **Pipeline threads** check `is_stale()` and the `discarding` flag to determine whether to drop their output.
- **Queue flushing** removes buffered audio and text, while `transport.discard_pending_audio()` drops pending client-side audio.
- **Response finalization** emits `response.done` with `status="cancelled"` and re-enables VAD listening via the `should_listen` event.
- **Generation isolation** ensures that late-arriving output from cancelled responses cannot contaminate new responses.

## Frequently Asked Questions

### What happens to audio already buffered when response.cancel is invoked?

The system clears the `output_queue` and `text_output_queue` via `_flush_queue()`, preserving only sentinel events that keep the pipeline alive. Additionally, `transport.discard_pending_audio()` drops any audio frames already transmitted to the client but not yet played.

### How does the generation counter prevent race conditions?

Each pipeline thread captures the current generation ID at the start of processing. When `cancel()` increments the global counter, threads detect the mismatch via `is_stale()` and silently discard their output. This prevents "ghost" audio from cancelled generations from reaching the client.

### Why does the system re-enable listening immediately after cancellation?

The `handle_response_cancel` method sets the `should_listen` event, which unblocks the VAD (Voice Activity Detection) component. This allows the user to speak again immediately without waiting for the cancelled generation to finish processing.

### Can a new response start while the previous one is still being cancelled?

Yes. When `response.create` arrives, it calls `cancel_scope.new_response()`, which clears the `discarding` flag and prepares the system for the new generation. The generation counter ensures that any lingering output from the previous response is ignored while the new response proceeds normally.