Speech-to-Speech Queue Types: How AudioInItem, TTSInItem, and Typed Queues Connect Pipeline Handlers
The speech-to-speech library uses eight strongly-typed queue aliases defined in queue_types.py—including AudioInItem, VADOutItem, STTOutItem, TextPromptItem, LMOutItem, TTSInItem, AudioOutItem, and TextEventItem—to enforce type-safe data flow between handlers, with each stage receiving a Queue[<InItem>] and emitting to a Queue[<OutItem>].
The Hugging Face speech-to-speech repository implements a modular real-time voice processing pipeline where data moves through consecutive stages via Python's Queue objects. Unlike generic byte streams, the system employs typed queues that carry specific payload objects, making the data flow explicit and statically verifiable. This architecture centers on src/speech_to_speech/pipeline/queue_types.py, which defines type aliases that determine exactly what data travels between the VAD, STT, LLM, and TTS handlers.
The Eight Queue Types Defined in queue_types.py
The pipeline's type safety relies on eight distinct TypeAlias definitions in src/speech_to_speech/pipeline/queue_types.py. Each alias represents the exact payload shape a handler expects on its input and produces on its output.
-
AudioInItem– CarriesVADInpayloads (raw audio bytes) entering fromSocketReceiver,WebSocketStreamer, orLocalAudioStreamer. Consumed exclusively byVADHandler. -
VADOutItem– CarriesVADAudioobjects (voice activity segments). Produced byVADHandlerand consumed by STT handlers likeWhisperSTTHandler. -
STTOutItem– CarriesPartialTranscriptionorTranscriptionobjects. Produced by any STT handler and consumed byTranscriptionNotifier. -
TextPromptItem– CarriesGenerateResponseRequestobjects. Produced byTranscriptionNotifierand consumed by LLM handlers includingResponsesApiModelHandler,ChatCompletionsApiModelHandler, andLanguageModelHandler. -
LMOutItem– CarriesLLMResponseChunk,TokenUsage, orEndOfResponseobjects. Produced by LLM handlers and consumed byLMOutputProcessor. -
TTSInItem– CarriesTTSInputorEndOfResponseobjects. Produced byLMOutputProcessorand consumed by TTS handlers likeChatTTSHandlerandQwen3TTSHandler. -
AudioOutItem– Carriesbytes,np.ndarray, orAudioOutputobjects. Produced by TTS handlers and consumed by output sinks likeLocalAudioStreamer,WebSocketStreamer, andSocketSender. -
TextEventItem– CarriesPipelineEventor sentinelbytes. Produced byLMOutputProcessoras a side-channel for real-time text updates and consumed by WebSocket or realtime clients.
This strongly typed system ensures that a mismatch between connected handlers would be caught by static type checkers before runtime.
How Handlers Are Wired in s2s_pipeline.py
The actual connection logic resides in _build_pipeline_handlers inside src/speech_to_speech/s2s_pipeline.py. Each handler receives specific queue instances created by initialize_queues_and_events (lines 34–48).
# From s2s_pipeline.py – handler instantiation showing queue wiring
vad = VADHandler(
stop_event,
queue_in=recv_audio_chunks_queue, # Queue[AudioInItem]
queue_out=spoken_prompt_queue, # Queue[VADOutItem]
)
transcription_notifier = TranscriptionNotifier(
stop_event,
queue_in=stt_output_queue, # Queue[STTOutItem]
queue_out=text_prompt_queue, # Queue[TextPromptItem]
)
# STT handler receives VADOutItem and emits STTOutItem
stt = get_stt_handler(..., spoken_prompt_queue, stt_output_queue, ...)
# LLM handler consumes TextPromptItem and produces LMOutItem
lm = get_llm_handler(..., text_prompt_queue, lm_response_queue, ...)
# LMOutputProcessor bridges LMOutItem to TTSInItem
lm_processor = LMOutputProcessor(
stop_event,
queue_in=lm_response_queue, # Queue[LMOutItem]
queue_out=lm_processed_queue, # Queue[TTSInItem]
setup_kwargs={"text_output_queue": text_output_queue}, # Queue[TextEventItem]
)
# TTS handler consumes TTSInItem and pushes AudioOutItem
tts = get_tts_handler(..., lm_processed_queue, send_audio_chunks_queue, ...)
This wiring creates a deterministic data flow: AudioInItem → VADOutItem → STTOutItem → TextPromptItem → LMOutItem → TTSInItem → AudioOutItem, with TextEventItem serving as a side channel for real-time text updates.
Practical Pipeline Construction Example
Below is a runnable example demonstrating how to instantiate the typed queues and wire handlers manually. This mirrors the internal logic in s2s_pipeline.py but isolates the components for clarity:
from queue import Queue
from threading import Event
from speech_to_speech.pipeline.queue_types import (
AudioInItem, VADOutItem, STTOutItem,
TextPromptItem, LMOutItem, TTSInItem,
AudioOutItem, TextEventItem,
)
from speech_to_speech.VAD.vad_handler import VADHandler
from speech_to_speech.STT.whisper_stt_handler import WhisperSTTHandler
from speech_to_speech.LLM.language_model import LanguageModelHandler
from speech_to_speech.LLM.lm_output_processor import LMOutputProcessor
from speech_to_speech.TTS.qwen3_tts_handler import Qwen3TTSHandler
from speech_to_speech.utils.thread_manager import ThreadManager
# Initialize typed queues
queues = {
"audio_in": Queue[AudioInItem](),
"vad_out": Queue[VADOutItem](),
"stt_out": Queue[STTOutItem](),
"text_prompt": Queue[TextPromptItem](),
"lm_out": Queue[LMOutItem](),
"tts_in": Queue[TTSInItem](),
"audio_out": Queue[AudioOutItem](),
"text_events": Queue[TextEventItem](),
}
stop = Event()
listen = Event()
# Wire handlers with correct queue types
vad = VADHandler(stop, queue_in=queues["audio_in"], queue_out=queues["vad_out"], setup_args=())
stt = WhisperSTTHandler(stop, queue_in=queues["vad_out"], queue_out=queues["stt_out"], setup_kwargs={})
lm = LanguageModelHandler(stop, queue_in=queues["text_prompt"], queue_out=queues["lm_out"], setup_kwargs={})
lm_proc = LMOutputProcessor(
stop,
queue_in=queues["lm_out"],
queue_out=queues["tts_in"],
setup_kwargs={"text_output_queue": queues["text_events"]}
)
tts = Qwen3TTSHandler(stop, queue_in=queues["tts_in"], queue_out=queues["audio_out"], setup_args=(listen,), setup_kwargs={})
# Execute pipeline
manager = ThreadManager([vad, stt, lm, lm_proc, tts])
manager.start()
# Feed raw audio into queues["audio_in"]...
# manager.stop() on completion
Each handler receives a Queue[<Item>] matching its declared type annotations, ensuring that VADHandler only processes AudioInItem payloads and produces VADOutItem objects for downstream consumers.
Key Source Files
Understanding the queue system requires familiarity with these specific files:
src/speech_to_speech/pipeline/queue_types.py– Central definition of all queue payload aliases (AudioInItem,VADOutItem, etc.)src/speech_to_speech/pipeline/handler_types.py– Generic handler signatures and protocol definitionssrc/speech_to_speech/pipeline/messages.py– Concrete Pydantic message classes (e.g.,VADAudio,Transcription,TTSInput)src/speech_to_speech/s2s_pipeline.py– Orchestrates queue creation (initialize_queues_and_events) and handler wiring (_build_pipeline_handlers)src/speech_to_speech/LLM/lm_output_processor.py– BridgesLMOutItemtoTTSInItemand emitsTextEventItemside channels
Summary
- Eight typed queues form the pipeline's backbone:
AudioInItem,VADOutItem,STTOutItem,TextPromptItem,LMOutItem,TTSInItem,AudioOutItem, andTextEventItem. - Type safety is enforced through
TypeAliasdefinitions inqueue_types.py, ensuring handlers receive exactly the data structure they expect. - Handler wiring occurs in
s2s_pipeline.pyvia_build_pipeline_handlers, where each component receivesQueue[<InItem>]andQueue[<OutItem>]instances. - Data flow follows a strict sequence: raw audio enters via
AudioInItem, flows through VAD/STT/LLM processing, and exits asAudioOutItemafter TTS synthesis. - Extensibility is simplified because new handlers only need to declare compatible input/output queue types to integrate with existing pipelines.
Frequently Asked Questions
What is the difference between AudioInItem and AudioOutItem?
AudioInItem carries VADIn payloads (raw audio bytes) entering the pipeline from sources like SocketReceiver or LocalAudioStreamer, and is consumed exclusively by VADHandler. AudioOutItem carries processed audio data (bytes, np.ndarray, or AudioOutput objects) leaving the TTS stage, consumed by output sinks like SocketSender or playback streams.
How does LMOutputProcessor bridge the LLM and TTS stages?
LMOutputProcessor (in src/speech_to_speech/LLM/lm_output_processor.py) consumes LMOutItem from the LLM handler, which may contain LLMResponseChunk, TokenUsage, or EndOfResponse objects. It processes these into TTSInput instances wrapped as TTSInItem, pushing them to the TTS queue. Additionally, it emits TextEventItem objects to a side-channel queue for real-time text streaming to clients.
Where are the queue types defined in the source code?
All queue type aliases are defined in src/speech_to_speech/pipeline/queue_types.py as TypeAlias assignments. The concrete message classes that populate these queues (like VADAudio or GenerateResponseRequest) are defined in src/speech_to_speech/pipeline/messages.py.
Why does the pipeline use typed queues instead of raw byte streams?
Typed queues provide compile-time safety and explicit data contracts between stages. Because each handler declares Queue[AudioInItem] or Queue[TTSInItem] in its constructor, static type checkers can verify that VADHandler never receives TTS output accidentally, and developers can trace the exact data shape flowing through the system by examining the TypeAlias definitions.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →