# Understanding TranscriptionNotifier: How It Enables Live Transcription Events in Hugging Face Speech-to-Speech

> Discover TranscriptionNotifier and how it powers live speech-to-text transcription events in Hugging Face. Learn about its real-time event emission and execution modes.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: how-to-guide
- Published: 2026-07-10

---

**TranscriptionNotifier is a `BaseHandler` that bridges speech-to-text (STT) and language model (LLM) components, emitting real-time `PartialTranscriptionEvent` and `TranscriptionCompletedEvent` objects to a thread-safe queue to enable live streaming of transcription results while supporting both realtime and legacy execution modes.**

The `huggingface/speech-to-speech` repository provides a modular pipeline for real-time voice AI systems. At the center of its live transcription capabilities sits the **TranscriptionNotifier**, a critical handler that transforms raw STT output into consumable events. This component ensures that partial transcriptions stream to clients in real-time while final utterances trigger downstream language model processing.

## Core Architecture and Pipeline Position

### Inheritance and Location

`TranscriptionNotifier` inherits from `BaseHandler` and is implemented in [`src/speech_to_speech/STT/transcription_notifier.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/STT/transcription_notifier.py). It sits between the STT component (which produces `PartialTranscription` and `Transcription` objects) and the LLM component (which consumes `GenerateResponseRequest` objects), acting as a translation layer that converts internal messages into public events.

### Event Types and Queue Management

The handler produces two distinct event types defined in [`src/speech_to_speech/pipeline/events.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/pipeline/events.py):
- **PartialTranscriptionEvent**: Carries incremental text updates for streaming display
- **TranscriptionCompletedEvent**: Signals final transcription completion with metadata including language code and timing

These events are placed on the `text_output_queue`—a thread-safe `Queue` instance established during the handler's `setup()` method.

## How TranscriptionNotifier Emits Live Transcription Events

### Streaming Partial Transcriptions

When the STT engine produces a `PartialTranscription` object, the notifier immediately creates a `PartialTranscriptionEvent` and puts it on the output queue. According to lines 44-51 in [`src/speech_to_speech/STT/transcription_notifier.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/STT/transcription_notifier.py), this occurs as soon as interim text arrives, allowing WebSocket layers or other downstream consumers to stream incremental updates to clients without waiting for the utterance to complete.

```python

# From the STT component

partial = PartialTranscription(text="Hello, ", turn_id=1, turn_revision=0)
list(notifier.process(partial))  # Enqueues PartialTranscriptionEvent, returns empty generator

```

### Signaling Final Transcriptions

Upon receiving a final `Transcription` object (lines 72-81), the handler queues a `TranscriptionCompletedEvent` containing the complete transcript, `language_code`, and `speech_stopped_at_s` metadata. This event signals that the speech segment has ended and the system can proceed to response generation.

## Dual Execution Mode Support

The handler adapts its behavior based on the presence of `runtime_config`, enabling both modern realtime and legacy request-response workflows.

### Realtime Mode (Event-Only Forwarding)

In realtime deployments where `runtime_config` is omitted, `TranscriptionNotifier` strictly limits itself to event emission. It pushes `PartialTranscriptionEvent` and `TranscriptionCompletedEvent` to the queue but yields no `GenerateResponseRequest`. The surrounding `RealtimeService` monitors these events and constructs generation requests independently, as demonstrated in [`tests/test_parakeet_transcription_events.py`](https://github.com/huggingface/speech-to-speech/blob/main/tests/test_parakeet_transcription_events.py).

### Legacy Mode (Direct LLM Bridging)

When `runtime_config` is provided, the handler operates in compatibility mode. After queueing the completed event, it appends the transcript to the chat history via `runtime_config.chat.add_item` and yields a `GenerateResponseRequest` (lines 95-103). This ensures the LLM handler receives uniform input regardless of whether the pipeline operates in realtime or batch mode.

## Edge Case Handling and Reliability

### Empty Transcript Protection

The handler includes safeguards against empty final transcriptions (lines 83-88). When a `Transcription` arrives with no text, the system still enqueues a `TranscriptionCompletedEvent` to maintain protocol consistency, but skips LLM request generation. If a `should_listen` flag is provided, the handler resets it via `should_listen.set()` to resume listening immediately (lines 85-87).

### Observability and Debugging

Comprehensive logging at each step (lines 52, 90-93) aids debugging by recording when partial results arrive, when final transcriptions complete, and when mode-specific branching occurs between realtime and legacy paths.

## Practical Implementation Example

The following example demonstrates how to instantiate the notifier, feed it partial and final transcriptions, and consume the resulting events:

```python
from queue import Queue
from threading import Event
from speech_to_speech.STT.transcription_notifier import TranscriptionNotifier
from speech_to_speech.pipeline.events import PartialTranscriptionEvent, TranscriptionCompletedEvent
from speech_to_speech.pipeline.messages import PartialTranscription, Transcription

# 1️⃣ Set up the notifier

output_q = Queue()
listen_flag = Event()
notifier = TranscriptionNotifier()
notifier.setup(text_output_queue=output_q, should_listen=listen_flag)

# 2️⃣ Feed a partial transcription (simulating an STT stream)

partial = PartialTranscription(text="Hello, ", turn_id=1, turn_revision=0)
list(notifier.process(partial))   # returns nothing, but enqueues an event

# 3️⃣ Consume the queued event

event = output_q.get()
assert isinstance(event, PartialTranscriptionEvent)
print(event.delta)   # → "Hello, "

# 4️⃣ Feed the final transcription

final = Transcription(text="Hello, world!", language_code="en", turn_id=1,
                     turn_revision=0, speech_stopped_at_s=2.3)
list(notifier.process(final))   # yields a GenerateResponseRequest in legacy mode

# 5️⃣ Retrieve the completed event

completed = output_q.get()
assert isinstance(completed, TranscriptionCompletedEvent)
print(completed.transcript)   # → "Hello, world!"

```

*In realtime deployments the `runtime_config` argument is omitted; the handler only pushes the two queue events, and the surrounding service builds a `GenerateResponseRequest` itself.*

## Summary

- `TranscriptionNotifier` inherits from `BaseHandler` and resides in [`src/speech_to_speech/STT/transcription_notifier.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/STT/transcription_notifier.py)
- It converts `PartialTranscription` objects into `PartialTranscriptionEvent` for immediate streaming (lines 44-51)
- It converts final `Transcription` objects into `TranscriptionCompletedEvent` to signal completion (lines 72-81)
- The handler supports **Realtime mode** (event-only) and **Legacy mode** (event + direct LLM request generation)
- Thread-safe queue management enables concurrent access by WebSocket layers and pipeline components
- Edge cases like empty transcripts are handled gracefully while maintaining the `should_listen` state

## Frequently Asked Questions

### What is the difference between PartialTranscriptionEvent and TranscriptionCompletedEvent?

`PartialTranscriptionEvent` carries incremental text updates via its `delta` attribute and is emitted continuously as the user speaks, enabling live transcription streaming. `TranscriptionCompletedEvent` is emitted once per utterance when the STT determines speech has stopped, containing the full `transcript`, `language_code`, and timing metadata to trigger downstream processing.

### When should I use Realtime mode versus Legacy mode with TranscriptionNotifier?

Use **Realtime mode** (omit `runtime_config`) when building low-latency applications where a separate service (like `RealtimeService`) manages response generation by monitoring the event queue. Use **Legacy mode** (provide `runtime_config`) when you need the handler to directly yield `GenerateResponseRequest` objects and maintain chat history, typically in batch processing or backward-compatible pipelines.

### How does TranscriptionNotifier ensure thread safety for live transcription events?

The handler relies on Python's standard `Queue` class (thread-safe by design) for event distribution. During setup, callers provide a `text_output_queue` that multiple threads can access safely. The `process()` method puts events onto this queue without blocking, allowing STT threads to emit events while WebSocket or LLM threads consume them concurrently.

### What happens if the STT returns an empty final transcription?

According to lines 83-88 in [`transcription_notifier.py`](https://github.com/huggingface/speech-to-speech/blob/main/transcription_notifier.py), the handler detects empty transcripts and skips LLM request generation to avoid processing silence. However, it still enqueues a `TranscriptionCompletedEvent` to maintain protocol consistency. If a `should_listen` Event object was provided during setup, the handler sets it to `True` (lines 85-87) to immediately resume listening for new speech.