Understanding TranscriptionNotifier: How It Enables Live Transcription Events in Hugging Face Speech-to-Speech
TranscriptionNotifier is a BaseHandler that bridges speech-to-text (STT) and language model (LLM) components, emitting real-time PartialTranscriptionEvent and TranscriptionCompletedEvent objects to a thread-safe queue to enable live streaming of transcription results while supporting both realtime and legacy execution modes.
The huggingface/speech-to-speech repository provides a modular pipeline for real-time voice AI systems. At the center of its live transcription capabilities sits the TranscriptionNotifier, a critical handler that transforms raw STT output into consumable events. This component ensures that partial transcriptions stream to clients in real-time while final utterances trigger downstream language model processing.
Core Architecture and Pipeline Position
Inheritance and Location
TranscriptionNotifier inherits from BaseHandler and is implemented in src/speech_to_speech/STT/transcription_notifier.py. It sits between the STT component (which produces PartialTranscription and Transcription objects) and the LLM component (which consumes GenerateResponseRequest objects), acting as a translation layer that converts internal messages into public events.
Event Types and Queue Management
The handler produces two distinct event types defined in src/speech_to_speech/pipeline/events.py:
- PartialTranscriptionEvent: Carries incremental text updates for streaming display
- TranscriptionCompletedEvent: Signals final transcription completion with metadata including language code and timing
These events are placed on the text_output_queue—a thread-safe Queue instance established during the handler's setup() method.
How TranscriptionNotifier Emits Live Transcription Events
Streaming Partial Transcriptions
When the STT engine produces a PartialTranscription object, the notifier immediately creates a PartialTranscriptionEvent and puts it on the output queue. According to lines 44-51 in src/speech_to_speech/STT/transcription_notifier.py, this occurs as soon as interim text arrives, allowing WebSocket layers or other downstream consumers to stream incremental updates to clients without waiting for the utterance to complete.
# From the STT component
partial = PartialTranscription(text="Hello, ", turn_id=1, turn_revision=0)
list(notifier.process(partial)) # Enqueues PartialTranscriptionEvent, returns empty generator
Signaling Final Transcriptions
Upon receiving a final Transcription object (lines 72-81), the handler queues a TranscriptionCompletedEvent containing the complete transcript, language_code, and speech_stopped_at_s metadata. This event signals that the speech segment has ended and the system can proceed to response generation.
Dual Execution Mode Support
The handler adapts its behavior based on the presence of runtime_config, enabling both modern realtime and legacy request-response workflows.
Realtime Mode (Event-Only Forwarding)
In realtime deployments where runtime_config is omitted, TranscriptionNotifier strictly limits itself to event emission. It pushes PartialTranscriptionEvent and TranscriptionCompletedEvent to the queue but yields no GenerateResponseRequest. The surrounding RealtimeService monitors these events and constructs generation requests independently, as demonstrated in tests/test_parakeet_transcription_events.py.
Legacy Mode (Direct LLM Bridging)
When runtime_config is provided, the handler operates in compatibility mode. After queueing the completed event, it appends the transcript to the chat history via runtime_config.chat.add_item and yields a GenerateResponseRequest (lines 95-103). This ensures the LLM handler receives uniform input regardless of whether the pipeline operates in realtime or batch mode.
Edge Case Handling and Reliability
Empty Transcript Protection
The handler includes safeguards against empty final transcriptions (lines 83-88). When a Transcription arrives with no text, the system still enqueues a TranscriptionCompletedEvent to maintain protocol consistency, but skips LLM request generation. If a should_listen flag is provided, the handler resets it via should_listen.set() to resume listening immediately (lines 85-87).
Observability and Debugging
Comprehensive logging at each step (lines 52, 90-93) aids debugging by recording when partial results arrive, when final transcriptions complete, and when mode-specific branching occurs between realtime and legacy paths.
Practical Implementation Example
The following example demonstrates how to instantiate the notifier, feed it partial and final transcriptions, and consume the resulting events:
from queue import Queue
from threading import Event
from speech_to_speech.STT.transcription_notifier import TranscriptionNotifier
from speech_to_speech.pipeline.events import PartialTranscriptionEvent, TranscriptionCompletedEvent
from speech_to_speech.pipeline.messages import PartialTranscription, Transcription
# 1️⃣ Set up the notifier
output_q = Queue()
listen_flag = Event()
notifier = TranscriptionNotifier()
notifier.setup(text_output_queue=output_q, should_listen=listen_flag)
# 2️⃣ Feed a partial transcription (simulating an STT stream)
partial = PartialTranscription(text="Hello, ", turn_id=1, turn_revision=0)
list(notifier.process(partial)) # returns nothing, but enqueues an event
# 3️⃣ Consume the queued event
event = output_q.get()
assert isinstance(event, PartialTranscriptionEvent)
print(event.delta) # → "Hello, "
# 4️⃣ Feed the final transcription
final = Transcription(text="Hello, world!", language_code="en", turn_id=1,
turn_revision=0, speech_stopped_at_s=2.3)
list(notifier.process(final)) # yields a GenerateResponseRequest in legacy mode
# 5️⃣ Retrieve the completed event
completed = output_q.get()
assert isinstance(completed, TranscriptionCompletedEvent)
print(completed.transcript) # → "Hello, world!"
In realtime deployments the runtime_config argument is omitted; the handler only pushes the two queue events, and the surrounding service builds a GenerateResponseRequest itself.
Summary
TranscriptionNotifierinherits fromBaseHandlerand resides insrc/speech_to_speech/STT/transcription_notifier.py- It converts
PartialTranscriptionobjects intoPartialTranscriptionEventfor immediate streaming (lines 44-51) - It converts final
Transcriptionobjects intoTranscriptionCompletedEventto signal completion (lines 72-81) - The handler supports Realtime mode (event-only) and Legacy mode (event + direct LLM request generation)
- Thread-safe queue management enables concurrent access by WebSocket layers and pipeline components
- Edge cases like empty transcripts are handled gracefully while maintaining the
should_listenstate
Frequently Asked Questions
What is the difference between PartialTranscriptionEvent and TranscriptionCompletedEvent?
PartialTranscriptionEvent carries incremental text updates via its delta attribute and is emitted continuously as the user speaks, enabling live transcription streaming. TranscriptionCompletedEvent is emitted once per utterance when the STT determines speech has stopped, containing the full transcript, language_code, and timing metadata to trigger downstream processing.
When should I use Realtime mode versus Legacy mode with TranscriptionNotifier?
Use Realtime mode (omit runtime_config) when building low-latency applications where a separate service (like RealtimeService) manages response generation by monitoring the event queue. Use Legacy mode (provide runtime_config) when you need the handler to directly yield GenerateResponseRequest objects and maintain chat history, typically in batch processing or backward-compatible pipelines.
How does TranscriptionNotifier ensure thread safety for live transcription events?
The handler relies on Python's standard Queue class (thread-safe by design) for event distribution. During setup, callers provide a text_output_queue that multiple threads can access safely. The process() method puts events onto this queue without blocking, allowing STT threads to emit events while WebSocket or LLM threads consume them concurrently.
What happens if the STT returns an empty final transcription?
According to lines 83-88 in transcription_notifier.py, the handler detects empty transcripts and skips LLM request generation to avoid processing silence. However, it still enqueues a TranscriptionCompletedEvent to maintain protocol consistency. If a should_listen Event object was provided during setup, the handler sets it to True (lines 85-87) to immediately resume listening for new speech.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →