How TranscriptionNotifier Taps Into Transcripts for Events in Speech-to-Speech Pipelines

The TranscriptionNotifier class in the huggingface/speech-to-speech repository intercepts speech-to-text output, converting PartialTranscription and Transcription objects into PartialTranscriptionEvent and TranscriptionCompletedEvent instances that it pushes onto the text_output_queue for WebSocket routing, while optionally yielding GenerateResponseRequest objects to trigger LLM processing in batch mode.

The TranscriptionNotifier serves as the critical bridge between automatic speech recognition and downstream processing components. By transforming raw STT output into typed pipeline events defined in src/speech_to_speech/pipeline/events.py, this handler enables both live transcription streaming to clients and structured handoff to language model handlers, making it essential for understanding the repository's event-driven architecture.

Where TranscriptionNotifier Sits in the Pipeline

In src/speech_to_speech/STT/transcription_notifier.py, the TranscriptionNotifier class implements the handler interface that consumes STTOut type aliases defined in src/speech_to_speech/pipeline/handler_types.py. This type union includes PartialTranscription and Transcription objects from src/speech_to_speech/pipeline/messages.py. The notifier sits between the STT handler and the LLM handler, acting as a transformation layer that prepares transcript data for both WebSocket distribution and language model consumption.

Processing Transcript Events

The notifier distinguishes between partial updates and final results to optimize the real-time experience while maintaining protocol consistency.

Handling Partial Transcriptions

When the process() method receives a PartialTranscription object, it immediately creates a PartialTranscriptionEvent containing the text delta and metadata such as turn_id and turn_revision. This event is placed on the text_output_queue—a queue consumed by the WebSocketRouter and RealtimeService dispatcher—to provide clients with "live" transcription feedback without waiting for final results.

Handling Final Transcriptions

For complete Transcription objects (or plain strings), the notifier builds a TranscriptionCompletedEvent containing the final transcript, optional language_code, identifiers (turn_id, turn_revision), and the timestamp when speech stopped (speech_stopped_at_s). If the transcript is empty, the notifier re-enables listening via the should_listen event and returns early. For non-empty transcripts, it logs the result and, when a runtime_config is present, appends the transcript to runtime_config.chat as a user message before yielding a GenerateResponseRequest for the LLM handler.

Setting Up the TranscriptionNotifier

Configure the notifier by providing the text_output_queue and optional runtime_config and should_listen flags via the setup() method:

from queue import Queue
from threading import Event
from speech_to_speech.STT.transcription_notifier import TranscriptionNotifier
from speech_to_speech.pipeline.events import PartialTranscriptionEvent, TranscriptionCompletedEvent

# Queue that will be consumed by the WebSocket router

text_queue = Queue()

# Optional flag to re‑enable listening after an empty transcript

listen_flag = Event()

notifier = TranscriptionNotifier()
notifier.setup(
    text_output_queue=text_queue,
    runtime_config=None,          # realtime mode

    should_listen=listen_flag,
)

Routing Partial Events to the WebSocket Layer

Process intermediate recognition results to stream deltas to connected clients:

from speech_to_speech.pipeline.messages import PartialTranscription

partial = PartialTranscription(
    text="Hel",
    turn_id="turn‑1",
    turn_revision=0,
)

# The notifier emits a PartialTranscriptionEvent onto the queue

list(notifier.process(partial))

event = text_queue.get()   # ← instance of PartialTranscriptionEvent

assert isinstance(event, PartialTranscriptionEvent)
assert event.delta == "Hel"

Real-Time vs. Batch Processing Logic

The notifier adapts its behavior based on whether runtime_config is provided, ensuring the LLM handler receives a uniform interface regardless of pipeline mode.

Real-Time Mode

When runtime_config is None, the notifier operates in real-time mode. It only forwards events to the text_output_queue and does not yield requests. This allows the pipeline to stream transcription updates to connected clients via the WebSocketRouter without triggering LLM generation.

Non-Real-Time Mode

When runtime_config is supplied (batch mode), the notifier updates runtime_config.chat with the transcript as a user message and yields a GenerateResponseRequest. This ensures the LLM handler receives a uniform request regardless of whether the pipeline is running in real-time or batch configuration.

from speech_to_speech.pipeline.messages import Transcription
from speech_to_speech.api.openai_realtime.runtime_config import RuntimeConfig

final = Transcription(
    text="Hello world",
    language_code="en",
    turn_id="turn‑1",
    turn_revision=0,
    speech_stopped_at_s=2.3,
)

runtime = RuntimeConfig()          # contains a chat object

notifier.setup(
    text_output_queue=text_queue,
    runtime_config=runtime,
    should_listen=None,
)

# The notifier yields a GenerateResponseRequest for the LLM handler

gen_requests = list(notifier.process(final))
assert len(gen_requests) == 1
request = gen_requests[0]

# The completed event is also placed on the queue

completed_event = text_queue.get()
assert isinstance(completed_event, TranscriptionCompletedEvent)
assert completed_event.transcript == "Hello world"

Integrating TranscriptionNotifier into the Pipeline

Wire the notifier between your STT and LLM handlers in the pipeline builder:


# In the pipeline builder (simplified)

pipeline = [
    WhisperSTTHandler(),
    TranscriptionNotifier(),
    LLMHandler(),
]

for handler in pipeline:
    handler.setup(...)

# Each handler yields items that the next handler consumes.

Summary

  • The TranscriptionNotifier in src/speech_to_speech/STT/transcription_notifier.py consumes STTOut types and produces typed events for downstream consumption.
  • It emits PartialTranscriptionEvent objects onto text_output_queue for live transcription updates and TranscriptionCompletedEvent objects for final results.
  • The text_output_queue serves as the communication channel to the WebSocketRouter and RealtimeService dispatcher.
  • In non-real-time mode, the notifier appends transcripts to runtime_config.chat and yields GenerateResponseRequest objects for LLM processing.
  • Empty transcripts trigger re-enabling of listening via should_listen without generating events or LLM requests.

Frequently Asked Questions

What is the difference between PartialTranscription and Transcription in the notifier?

PartialTranscription represents intermediate recognition results containing partial text deltas, while Transcription represents the final, complete utterance with additional metadata like language_code and speech_stopped_at_s. The notifier wraps partial inputs in PartialTranscriptionEvent for immediate streaming to clients, whereas final inputs become TranscriptionCompletedEvent objects that may also trigger LLM generation in batch mode.

How does TranscriptionNotifier handle empty transcripts?

When the notifier receives an empty transcript, it checks the should_listen event flag and re-enables listening before returning early without placing any event on the queue. This prevents the pipeline from processing silence or non-speech segments as valid input, ensuring that only meaningful transcripts trigger downstream processing or client notifications.

What queue does TranscriptionNotifier use to communicate with the client?

The notifier communicates with the client via the text_output_queue provided during setup, which is typically consumed by the WebSocketRouter or RealtimeService dispatcher. Both PartialTranscriptionEvent and TranscriptionCompletedEvent objects defined in src/speech_to_speech/pipeline/events.py are placed on this queue, maintaining a consistent event schema throughout the real-time protocol layer.

How does the notifier distinguish between real-time and batch processing?

The notifier distinguishes modes by checking whether runtime_config is None (real-time) or populated (batch). In real-time mode, it only emits events to the queue; in batch mode, it additionally appends the transcript to runtime_config.chat and yields a GenerateResponseRequest for LLM processing, ensuring the downstream handler receives a uniform interface regardless of pipeline configuration.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →