How response.cancel Interrupts Generation in the Speech-to-Speech Realtime Engine
The response.cancel mechanism aborts active LLM and TTS generation by incrementing a generation counter and setting a discarding flag, which pipeline threads consult to drop stale output and immediately restore the system to a listening state.
The huggingface/speech-to-speech repository implements a realtime audio pipeline that requires robust interruption capabilities. When users need to stop an ongoing response, the response.cancel event provides immediate termination of generation and flushes pending audio queues through a coordinated cancellation protocol.
The Cancellation Architecture
The interruption flow centers on three core components that work together to safely abort generation without corrupting pipeline state.
CancelScope State Management
The CancelScope class in src/speech_to_speech/pipeline/cancel_scope.py maintains a generation counter (_gen) and a discarding flag (_discarding). These properties allow pipeline threads to determine whether their output is still valid or belongs to a cancelled generation.
WebSocket Router Handling
The _dispatch_client_event function in src/speech_to_speech/api/openai_realtime/websocket_router.py receives the ResponseCancelEvent and triggers the cancellation sequence. It invokes unit.cancel_scope.cancel(), flushes pending queues, and forwards the event to the service layer.
ResponseHandler Finalization
The handle_response_cancel method in src/speech_to_speech/api/openai_realtime/handlers/response.py emits the response.done event with status="cancelled" and re-enables the should_listen event to restore voice activity detection.
Step-by-Step Cancellation Flow
1. Client Sends the Cancel Event
The client transmits a JSON message over the WebSocket:
{
"type": "response.cancel"
}
2. Router Triggers CancelScope
In websocket_router.py, the _dispatch_client_event function detects the ResponseCancelEvent and executes the following logic:
elif isinstance(event, ResponseCancelEvent):
was_active = service._state(session_id).in_response
if was_active:
unit.cancel_scope.cancel() # ← bump generation, set discarding
_flush_queue(unit.output_queue, preserve=_keep_audio_sentinel)
_flush_queue(unit.text_output_queue, preserve=_keep_user_text_event)
transport.discard_pending_audio()
events = service.handle_response_cancel(session_id)
await transport.send_events(events)
unit.response_playing.clear()
This sequence performs four critical actions:
unit.cancel_scope.cancel()increments the generation counter and sets_discarding = True- Queue flushing removes buffered audio and text belonging to the cancelled response
transport.discard_pending_audio()drops audio already sent to the client but not yet played- Event forwarding notifies the service layer to finalize the response
3. Pipeline Threads Check Generation Counters
The CancelScope class implements the core cancellation semantics:
class CancelScope:
def __init__(self):
self._gen = 0
self._discarding = False
self._discarded_generation = None
def cancel(self) -> None:
self._discarded_generation = self._gen
self._gen = (self._gen + 1) & 0xFFFFFFFF # wrap-around safety
self._discarding = True
def is_stale(self, gen: int) -> bool:
return gen != self._gen
@property
def discarding(self) -> bool:
return self._discarding
Each pipeline thread captures cancel_scope.generation at the start of processing. If cancel() has incremented the counter, is_stale() returns True, causing the thread to stop processing and drop its output.
4. Service Emits response.done
The ResponseHandler.handle_response_cancel method finalizes the cancelled response:
def handle_response_cancel(self, conn_id: str) -> list[ServerEvent]:
events = self.finish_response(conn_id, status="cancelled", reason="client_cancelled")
should_listen = self._should_listen(conn_id)
if should_listen:
should_listen.set() # re-enable VAD listening
logger.info("Response cancelled, listening re-enabled")
return events
This emits response.done with status="cancelled" and sets the should_listen event to unblock the VAD.
Implementation Details
Send-Loop Guard Mechanism
Inside the per-unit send loop (websocket_router._send_loop_for), the code guards against emitting stale audio:
if unit.cancel_scope.discarding and generation != unit.cancel_scope.generation:
# Discard stale audio while a cancel is in effect
continue
Once cancel() has been called, any audio generated with the old generation is ignored until a new response.create arrives.
Reset on New Response
When a fresh response.create event arrives, the router invokes:
unit.cancel_scope.new_response()
This clears the discarding flag and the discarded generation marker, allowing the next generation to flow normally.
Practical Code Examples
Server-Side Manual Cancellation
You can trigger cancellation programmatically from custom handlers:
from speech_to_speech.pipeline.cancel_scope import CancelScope
def my_handler(unit):
# Detect user interruption (e.g., hot-key)
if user_pressed_stop():
unit.cancel_scope.cancel() # abort current generation
# Push sentinel to unblock the send loop
unit.output_queue.put(AudioOutput(audio=AUDIO_RESPONSE_DONE,
cancel_generation=unit.cancel_scope.generation))
Client-Side JavaScript Implementation
// ws is an open WebSocket to /v1/realtime
function cancelCurrentResponse() {
ws.send(JSON.stringify({type: "response.cancel"}));
}
Testing the Cancel Flow
def test_response_cancel_flushes_queues(setup):
app, service, _, output_q, text_q, _, _, _, cancel_scope = setup
ws.send_json({"type": "response.cancel"})
assert cancel_scope.discarding # flag is set
# After the server sends response.done, the flag is cleared
await wait_for(lambda: not cancel_scope.discarding)
Key Source Files
| File | Purpose |
|---|---|
src/speech_to_speech/pipeline/cancel_scope.py |
Core cancellation state (generation counter, discarding flag) |
src/speech_to_speech/api/openai_realtime/websocket_router.py |
Entry point for response.cancel; triggers queue flushing |
src/speech_to_speech/api/openai_realtime/handlers/response.py |
Generates response.done with status="cancelled"; re-enables listening |
tests/openai_realtime/test_websocket_router.py |
Validates the cancel flow (discarding flag, queue flushing) |
Summary
response.cancelimmediately interrupts LLM and TTS generation by incrementing a generation counter inCancelScope.- Pipeline threads check
is_stale()and thediscardingflag to determine whether to drop their output. - Queue flushing removes buffered audio and text, while
transport.discard_pending_audio()drops pending client-side audio. - Response finalization emits
response.donewithstatus="cancelled"and re-enables VAD listening via theshould_listenevent. - Generation isolation ensures that late-arriving output from cancelled responses cannot contaminate new responses.
Frequently Asked Questions
What happens to audio already buffered when response.cancel is invoked?
The system clears the output_queue and text_output_queue via _flush_queue(), preserving only sentinel events that keep the pipeline alive. Additionally, transport.discard_pending_audio() drops any audio frames already transmitted to the client but not yet played.
How does the generation counter prevent race conditions?
Each pipeline thread captures the current generation ID at the start of processing. When cancel() increments the global counter, threads detect the mismatch via is_stale() and silently discard their output. This prevents "ghost" audio from cancelled generations from reaching the client.
Why does the system re-enable listening immediately after cancellation?
The handle_response_cancel method sets the should_listen event, which unblocks the VAD (Voice Activity Detection) component. This allows the user to speak again immediately without waiting for the cancelled generation to finish processing.
Can a new response start while the previous one is still being cancelled?
Yes. When response.create arrives, it calls cancel_scope.new_response(), which clears the discarding flag and prepares the system for the new generation. The generation counter ensures that any lingering output from the previous response is ignored while the new response proceeds normally.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →