# How CancelScope Enables Interruption of LLM and TTS Responses in the Speech-to-Speech Pipeline

> Discover how CancelScope instantly invalidates LLM text and TTS audio with a generation counter and discarding flag, enabling seamless interruption in speech-to-speech pipelines.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: internals
- Published: 2026-07-11

---

**CancelScope uses a generation counter and discarding flag to instantly invalidate in-flight LLM text and TTS audio when a user interrupts the speech-to-speech pipeline.**

The `huggingface/speech-to-speech` repository implements a real-time conversation system that chains asynchronous handlers (VAD → STT → LLM → TTS) to stream audio responses. When a user triggers a stop command, the pipeline must immediately halt language model generation and discard any buffered text-to-speech output without race conditions. This coordination is achieved through the **CancelScope** object, a lightweight synchronization primitive shared across all pipeline components.

## The Dual-Mechanism Design of CancelScope

CancelScope provides two complementary mechanisms to ensure clean interruption across distributed async handlers. According to the source code in [`src/speech_to_speech/pipeline/cancel_scope.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/pipeline/cancel_scope.py), the class maintains atomic state for generation tracking and output suppression.

### Generation Counter for Stale Detection

The core of the cancellation logic relies on a monotonic **generation counter** (`self._gen`). Each handler captures the current generation at the start of processing a response. When `cancel()` is invoked, the counter increments atomically, rendering any previously captured generation values stale.

Handlers verify validity by calling `CancelScope.is_stale(generation)`, which compares the captured generation against the current counter value. As implemented in [`src/speech_to_speech/baseHandler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/baseHandler.py), the `BaseHandler.handle()` method checks this state before forwarding chunks:

```python
gen = cancel_scope.generation

# ... process chunk ...

if cancel_scope.is_stale(cancel_generation):
    return  # drop stale output

```

### Discarding Flag for Audio Suppression

The **discarding flag** (`self._discarding`) solves the problem of buffered audio packets that were generated before cancellation but arrive at the output queue afterward. When `cancel()` is called, it sets `_discarding = True` and stores the discarded generation in `_discarded_generation`.

The WebSocket router in [`src/speech_to_speech/api/openai_realtime/websocket_router.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/api/openai_realtime/websocket_router.py) inspects this flag in its async send loop. Before transmitting audio to the client, it verifies that the packet's generation matches the current scope:

```python
if unit.cancel_scope.discarding and generation != unit.cancel_scope.generation:
    continue  # silently discard stale packet

```

Only when the final "response-done" marker arrives does the router call `response_done(generation)` to clear the discarding flag.

## Step-by-Step Cancellation Flow

The interruption workflow follows a strict sequence across the pipeline components:

1. **Pipeline Initialization** – A fresh `CancelScope` is instantiated and injected into the shared handler kwargs during construction in [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py):
   ```python
   cancel_scope = CancelScope()
   vars(kw)["cancel_scope"] = cancel_scope
   ```

2. **Handler Start** – Each handler receives the shared scope. At the beginning of a new LLM or TTS response, the handler captures the current generation to tag all subsequent output.

3. **User-Initiated Cancel** – The WebSocket router (or external controller) invokes `cancel_scope.cancel()`. This atomically increments the generation counter and sets the discarding flag to True.

4. **Stale-Check in Handlers** – While processing chunks, handlers consult `is_stale()`. If the generation no longer matches, the handler returns early, stopping further processing and preventing new output from entering the queue.

5. **Discard Guard in Send Loop** – Before any audio packet reaches the client, the router checks the discarding flag. Packets with mismatched generations are dropped silently, ensuring no "ghost" audio from the cancelled response plays.

## Implementation in the Pipeline Code

The following patterns demonstrate how CancelScope integrates with the actual handler implementations. When building a custom handler in the `huggingface/speech-to-speech` framework, you must respect the cancellation contract:

**Capturing generation in LLM handlers:**

```python
from speech_to_speech.pipeline.cancel_scope import CancelScope

def handle_llm_chunk(self, chunk):
    # Capture generation at response start

    gen = self.cancel_scope.generation
    
    # Tag output with generation for downstream filtering

    self.output_queue.put(AssistantTextEvent(
        text=chunk, 
        cancel_generation=gen
    ))

```

**Triggering cancellation on user interrupt:**

```python
def user_stop_requested():
    # Increments generation and sets discarding flag

    cancel_scope.cancel()

```

**Filtering stale audio in WebSocket transport:**

```python
async def send_loop(unit):
    while True:
        item = output_queue.get()
        gen = getattr(item, "cancel_generation", None)
        
        # Check discarding flag in CancelScope

        if unit.cancel_scope.discarding and gen != unit.cancel_scope.generation:
            continue
            
        await ws.send(item.audio)

```

The design leverages Python's GIL to ensure atomic reads and writes of the integer counter and boolean flag, eliminating the need for explicit locks while maintaining thread safety across async boundaries.

## Summary

- **CancelScope** coordinates interruption across the VAD → STT → LLM → TTS pipeline using a shared generation counter and discarding flag.
- The **generation counter** (`_gen`) invalidates in-flight work by incrementing when `cancel()` is called; handlers check `is_stale()` to abort processing.
- The **discarding flag** (`_discarding`) prevents stale audio packets from reaching the client when buffered output arrives after cancellation.
- Key implementations reside in [`src/speech_to_speech/pipeline/cancel_scope.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/pipeline/cancel_scope.py), [`src/speech_to_speech/baseHandler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/baseHandler.py), and the WebSocket router at [`src/speech_to_speech/api/openai_realtime/websocket_router.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/api/openai_realtime/websocket_router.py).
- The mechanism is lock-free, relying on Python's GIL for atomic state updates, making it suitable for high-throughput real-time audio streaming.

## Frequently Asked Questions

### What is CancelScope in the speech-to-speech pipeline?

CancelScope is a synchronization object that manages cancellation state across the distributed asynchronous handlers in the `huggingface/speech-to-speech` pipeline. It tracks generation numbers and discarding status to ensure that when a user interrupts the conversation, any in-progress LLM generation and TTS synthesis is immediately halted and discarded.

### How does CancelScope prevent race conditions during cancellation?

CancelScope prevents race conditions by using a monotonic generation counter and atomic boolean flag. When `cancel()` is called, it increments the counter and sets the discarding flag in a single operation. Since Python's GIL ensures atomic reads/writes for integers and booleans, handlers either see the pre-cancellation or post-cancellation state without intermediate values, eliminating torn reads or synchronization bugs.

### Where is CancelScope initialized in the codebase?

CancelScope is instantiated during pipeline construction in [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py). The object is created once per pipeline or realtime unit and injected into the shared kwargs dictionary (`vars(kw)["cancel_scope"]`), ensuring every handler receives the same scope instance via `handler_kwargs`.

### Can CancelScope handle multiple concurrent interruptions?

Yes, CancelScope handles rapid successive cancellations correctly. Each call to `cancel()` increments the generation counter and updates the discarded generation marker. If a second cancellation occurs before the first response fully clears, the generation counter advances again, and the discarding flag remains active until the router processes the final response-done marker for the latest generation.