What Is the TranscriptionNotifier and How Does It Manage Text Output Queues in Speech-to-Speech?
The TranscriptionNotifier is a pipeline handler that bridges Speech-to-Text (STT) and Large Language Model (LLM) components, managing text output queues by emitting partial and final transcription events while optionally forwarding requests to the LLM in legacy mode.
The TranscriptionNotifier serves as a critical intermediary in the huggingface/speech-to-speech repository, orchestrating the flow of transcription data between speech recognition and language generation. This handler manages thread-safe text output queues to deliver real-time updates to WebSocket clients while maintaining compatibility with legacy pipeline configurations. Understanding how the TranscriptionNotifier handles both partial and final transcriptions is essential for building robust speech-to-speech applications.
Core Architecture and Operating Modes
The handler operates in two distinct modes depending on the presence of a runtime_config object.
Realtime Mode (No Runtime Configuration)
When initialized without a runtime_config, the notifier functions as a pure event emitter. It pushes partial and final transcription events onto the text_output_queue for consumption by the WebSocket protocol layer, but does not forward any data to the LLM. This mode is ideal for real-time transcription display where immediate language generation is not required.
Legacy Mode (With Runtime Configuration)
When a runtime_config is provided, the notifier performs additional processing beyond queue management. It appends user messages to the RuntimeConfig.chat object and yields GenerateResponseRequest instances, ensuring the downstream LLM handler receives uniform input regardless of execution mode. This maintains backward compatibility with existing pipeline implementations.
Internal Queue Management Mechanisms
The notifier's queue management follows a strict lifecycle defined in src/speech_to_speech/STT/transcription_notifier.py.
Setup and Initialization
The setup method receives three optional parameters that define the handler's behavior:
def setup(self, text_output_queue=None, runtime_config=None, should_listen=None):
self.text_output_queue = text_output_queue
self.runtime_config = runtime_config
self.should_listen = should_listen
text_output_queue: AQueue[TextEventItem]where pipeline events are pushed.runtime_config: An optionalRuntimeConfigused by legacy pipelines.should_listen: AnEventindicating whether the system should resume listening after an empty final transcription.
Handling Partial Transcriptions
When processing a PartialTranscription, the notifier immediately places a PartialTranscriptionEvent onto the queue if the queue exists and the transcription contains text. It then returns nothing, keeping the downstream LLM untouched while informing the client of live text updates.
if isinstance(transcription, PartialTranscription):
if self.text_output_queue and transcription.text:
self.text_output_queue.put(
PartialTranscriptionEvent(delta=str(transcription.text))
)
return
This approach ensures that interim speech recognition results are visible to users without triggering premature LLM processing.
Processing Final Transcriptions
For completed Transcription objects (or raw strings), the notifier executes a more complex workflow:
- Constructs a
TranscriptionCompletedEventand pushes it onto the queue. - Logs the transcript with language code if present.
- Handles empty transcripts by re-enabling listening via
should_listen.set(). - For non-empty transcripts with a
runtime_config, adds the text as a user message and yields aGenerateResponseRequest.
if self.text_output_queue is not None:
self.text_output_queue.put(TranscriptionCompletedEvent(...))
if not transcript:
if self.should_listen is not None:
self.should_listen.set()
return
if self.runtime_config is not None:
self.runtime_config.chat.add_item(make_user_message(transcript))
yield GenerateResponseRequest(runtime_config=self.runtime_config)
This logic ensures that empty final transcriptions still trigger completion events for the client while avoiding unnecessary LLM requests.
Event Types and Queue Structure
The queue holds typed events defined in src/speech_to_speech/pipeline/events.py. The notifier pushes two primary event types:
PartialTranscriptionEvent: Carries live text deltas representing incomplete speech recognition results.TranscriptionCompletedEvent: Carries the final transcript, language code, turn identifiers, and timestamp (speech_stopped_at_s).
These events are consumed by the realtime WebSocket router (RealtimeService.dispatch_pipeline_event), which translates them into OpenAI Realtime protocol messages. By emitting both partial and completed events even for empty final transcriptions, the notifier guarantees that the client sees a consistent lifecycle of a speech turn.
Practical Implementation Examples
Realtime Mode Setup
The following example demonstrates configuring the notifier for realtime transcription without LLM integration:
from queue import Queue
from speech_to_speech.STT.transcription_notifier import TranscriptionNotifier
from speech_to_speech.pipeline.messages import PartialTranscription, Transcription
queue = Queue()
notifier = TranscriptionNotifier()
notifier.setup(text_output_queue=queue) # realtime mode, no runtime_config
# Feed a partial transcription
notifier.process(PartialTranscription(text="Hel"))
# → queue now contains a PartialTranscriptionEvent(delta="Hel")
# Feed the final transcription
notifier.process(Transcription(text="Hello", language_code="en"))
# → queue now contains a TranscriptionCompletedEvent(transcript="Hello")
Legacy Mode with LLM Integration
This example shows legacy mode configuration where transcriptions trigger LLM generation:
from threading import Event
from queue import Queue
from speech_to_speech.api.openai_realtime.runtime_config import RuntimeConfig
from speech_to_speech.STT.transcription_notifier import TranscriptionNotifier
from speech_to_speech.pipeline.messages import Transcription
queue = Queue()
runtime_cfg = RuntimeConfig()
listen_evt = Event()
notifier = TranscriptionNotifier()
notifier.setup(
text_output_queue=queue,
runtime_config=runtime_cfg,
should_listen=listen_evt,
)
# A non-empty final transcription triggers LLM generation
gen_requests = list(
notifier.process(Transcription(text="Hi there", language_code="en"))
)
assert len(gen_requests) == 1
assert gen_requests[0].runtime_config is runtime_cfg
Summary
- The TranscriptionNotifier acts as a boundary handler between STT and LLM components, supporting both realtime and legacy execution modes.
- Queue management relies on thread-safe
Queueobjects emittingPartialTranscriptionEventandTranscriptionCompletedEventinstances defined insrc/speech_to_speech/pipeline/events.py. - Empty transcript handling ensures the client receives completion signals via
should_listenevents while preventing unnecessary LLM requests. - Legacy mode integration appends messages to
RuntimeConfig.chatand yieldsGenerateResponseRequestobjects for downstream processing. - Source files include
src/speech_to_speech/STT/transcription_notifier.pyfor the core implementation andsrc/speech_to_speech/pipeline/messages.pyfor message type definitions.
Frequently Asked Questions
What is the primary purpose of the TranscriptionNotifier?
The TranscriptionNotifier serves as a pipeline handler that manages the transition between speech recognition and language generation. It encapsulates the logic for emitting transcription events to WebSocket clients while optionally preparing LLM requests in legacy mode, effectively isolating the STT-to-LLM boundary.
How does the TranscriptionNotifier handle empty transcriptions?
When receiving an empty final transcription, the notifier pushes a TranscriptionCompletedEvent to the queue to signal completion to the client, then sets the should_listen event to resume audio capture. It returns early without generating an LLM request, preventing the system from processing silence or non-speech audio.
What is the difference between PartialTranscriptionEvent and TranscriptionCompletedEvent?
PartialTranscriptionEvent carries incremental text deltas representing incomplete recognition results, allowing realtime display of speech as it is spoken. TranscriptionCompletedEvent contains the final transcript, language code, turn identifiers, and timestamp, marking the definitive end of a speech turn.
How does the notifier ensure thread safety when managing queues?
The handler uses Python's standard Queue class from the queue module, which provides thread-safe operations for putting and getting items. This allows safe communication between the synchronous STT processing thread and the asynchronous WebSocket send loop without requiring explicit locks.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →