How response.cancel Interrupts Generation in the Speech-to-Speech Realtime Engine

The response.cancel mechanism aborts active LLM and TTS generation by incrementing a generation counter and setting a discarding flag, which pipeline threads consult to drop stale output and immediately restore the system to a listening state.

The huggingface/speech-to-speech repository implements a realtime audio pipeline that requires robust interruption capabilities. When users need to stop an ongoing response, the response.cancel event provides immediate termination of generation and flushes pending audio queues through a coordinated cancellation protocol.

The Cancellation Architecture

The interruption flow centers on three core components that work together to safely abort generation without corrupting pipeline state.

CancelScope State Management

The CancelScope class in src/speech_to_speech/pipeline/cancel_scope.py maintains a generation counter (_gen) and a discarding flag (_discarding). These properties allow pipeline threads to determine whether their output is still valid or belongs to a cancelled generation.

WebSocket Router Handling

The _dispatch_client_event function in src/speech_to_speech/api/openai_realtime/websocket_router.py receives the ResponseCancelEvent and triggers the cancellation sequence. It invokes unit.cancel_scope.cancel(), flushes pending queues, and forwards the event to the service layer.

ResponseHandler Finalization

The handle_response_cancel method in src/speech_to_speech/api/openai_realtime/handlers/response.py emits the response.done event with status="cancelled" and re-enables the should_listen event to restore voice activity detection.

Step-by-Step Cancellation Flow

1. Client Sends the Cancel Event

The client transmits a JSON message over the WebSocket:

{
  "type": "response.cancel"
}

2. Router Triggers CancelScope

In websocket_router.py, the _dispatch_client_event function detects the ResponseCancelEvent and executes the following logic:

elif isinstance(event, ResponseCancelEvent):
    was_active = service._state(session_id).in_response
    if was_active:
        unit.cancel_scope.cancel()          # ← bump generation, set discarding

    _flush_queue(unit.output_queue, preserve=_keep_audio_sentinel)
    _flush_queue(unit.text_output_queue, preserve=_keep_user_text_event)
    transport.discard_pending_audio()
    events = service.handle_response_cancel(session_id)
    await transport.send_events(events)
    unit.response_playing.clear()

This sequence performs four critical actions:

  • unit.cancel_scope.cancel() increments the generation counter and sets _discarding = True
  • Queue flushing removes buffered audio and text belonging to the cancelled response
  • transport.discard_pending_audio() drops audio already sent to the client but not yet played
  • Event forwarding notifies the service layer to finalize the response

3. Pipeline Threads Check Generation Counters

The CancelScope class implements the core cancellation semantics:

class CancelScope:
    def __init__(self):
        self._gen = 0
        self._discarding = False
        self._discarded_generation = None

    def cancel(self) -> None:
        self._discarded_generation = self._gen
        self._gen = (self._gen + 1) & 0xFFFFFFFF   # wrap-around safety

        self._discarding = True

    def is_stale(self, gen: int) -> bool:
        return gen != self._gen

    @property
    def discarding(self) -> bool:
        return self._discarding

Each pipeline thread captures cancel_scope.generation at the start of processing. If cancel() has incremented the counter, is_stale() returns True, causing the thread to stop processing and drop its output.

4. Service Emits response.done

The ResponseHandler.handle_response_cancel method finalizes the cancelled response:

def handle_response_cancel(self, conn_id: str) -> list[ServerEvent]:
    events = self.finish_response(conn_id, status="cancelled", reason="client_cancelled")
    should_listen = self._should_listen(conn_id)
    if should_listen:
        should_listen.set()                     # re-enable VAD listening

    logger.info("Response cancelled, listening re-enabled")
    return events

This emits response.done with status="cancelled" and sets the should_listen event to unblock the VAD.

Implementation Details

Send-Loop Guard Mechanism

Inside the per-unit send loop (websocket_router._send_loop_for), the code guards against emitting stale audio:

if unit.cancel_scope.discarding and generation != unit.cancel_scope.generation:
    # Discard stale audio while a cancel is in effect

    continue

Once cancel() has been called, any audio generated with the old generation is ignored until a new response.create arrives.

Reset on New Response

When a fresh response.create event arrives, the router invokes:

unit.cancel_scope.new_response()

This clears the discarding flag and the discarded generation marker, allowing the next generation to flow normally.

Practical Code Examples

Server-Side Manual Cancellation

You can trigger cancellation programmatically from custom handlers:

from speech_to_speech.pipeline.cancel_scope import CancelScope

def my_handler(unit):
    # Detect user interruption (e.g., hot-key)

    if user_pressed_stop():
        unit.cancel_scope.cancel()          # abort current generation

        # Push sentinel to unblock the send loop

        unit.output_queue.put(AudioOutput(audio=AUDIO_RESPONSE_DONE,
                                          cancel_generation=unit.cancel_scope.generation))

Client-Side JavaScript Implementation

// ws is an open WebSocket to /v1/realtime
function cancelCurrentResponse() {
  ws.send(JSON.stringify({type: "response.cancel"}));
}

Testing the Cancel Flow

def test_response_cancel_flushes_queues(setup):
    app, service, _, output_q, text_q, _, _, _, cancel_scope = setup
    ws.send_json({"type": "response.cancel"})
    assert cancel_scope.discarding          # flag is set

    # After the server sends response.done, the flag is cleared

    await wait_for(lambda: not cancel_scope.discarding)

Key Source Files

File Purpose
src/speech_to_speech/pipeline/cancel_scope.py Core cancellation state (generation counter, discarding flag)
src/speech_to_speech/api/openai_realtime/websocket_router.py Entry point for response.cancel; triggers queue flushing
src/speech_to_speech/api/openai_realtime/handlers/response.py Generates response.done with status="cancelled"; re-enables listening
tests/openai_realtime/test_websocket_router.py Validates the cancel flow (discarding flag, queue flushing)

Summary

  • response.cancel immediately interrupts LLM and TTS generation by incrementing a generation counter in CancelScope.
  • Pipeline threads check is_stale() and the discarding flag to determine whether to drop their output.
  • Queue flushing removes buffered audio and text, while transport.discard_pending_audio() drops pending client-side audio.
  • Response finalization emits response.done with status="cancelled" and re-enables VAD listening via the should_listen event.
  • Generation isolation ensures that late-arriving output from cancelled responses cannot contaminate new responses.

Frequently Asked Questions

What happens to audio already buffered when response.cancel is invoked?

The system clears the output_queue and text_output_queue via _flush_queue(), preserving only sentinel events that keep the pipeline alive. Additionally, transport.discard_pending_audio() drops any audio frames already transmitted to the client but not yet played.

How does the generation counter prevent race conditions?

Each pipeline thread captures the current generation ID at the start of processing. When cancel() increments the global counter, threads detect the mismatch via is_stale() and silently discard their output. This prevents "ghost" audio from cancelled generations from reaching the client.

Why does the system re-enable listening immediately after cancellation?

The handle_response_cancel method sets the should_listen event, which unblocks the VAD (Voice Activity Detection) component. This allows the user to speak again immediately without waiting for the cancelled generation to finish processing.

Can a new response start while the previous one is still being cancelled?

Yes. When response.create arrives, it calls cancel_scope.new_response(), which clears the discarding flag and prepares the system for the new generation. The generation counter ensures that any lingering output from the previous response is ignored while the new response proceeds normally.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →