Speech-to-Speech Queue Types: How AudioInItem, TTSInItem, and Typed Queues Connect Pipeline Handlers

The speech-to-speech library uses eight strongly-typed queue aliases defined in queue_types.py—including AudioInItem, VADOutItem, STTOutItem, TextPromptItem, LMOutItem, TTSInItem, AudioOutItem, and TextEventItem—to enforce type-safe data flow between handlers, with each stage receiving a Queue[<InItem>] and emitting to a Queue[<OutItem>].

The Hugging Face speech-to-speech repository implements a modular real-time voice processing pipeline where data moves through consecutive stages via Python's Queue objects. Unlike generic byte streams, the system employs typed queues that carry specific payload objects, making the data flow explicit and statically verifiable. This architecture centers on src/speech_to_speech/pipeline/queue_types.py, which defines type aliases that determine exactly what data travels between the VAD, STT, LLM, and TTS handlers.

The Eight Queue Types Defined in queue_types.py

The pipeline's type safety relies on eight distinct TypeAlias definitions in src/speech_to_speech/pipeline/queue_types.py. Each alias represents the exact payload shape a handler expects on its input and produces on its output.

  • AudioInItem – Carries VADIn payloads (raw audio bytes) entering from SocketReceiver, WebSocketStreamer, or LocalAudioStreamer. Consumed exclusively by VADHandler.

  • VADOutItem – Carries VADAudio objects (voice activity segments). Produced by VADHandler and consumed by STT handlers like WhisperSTTHandler.

  • STTOutItem – Carries PartialTranscription or Transcription objects. Produced by any STT handler and consumed by TranscriptionNotifier.

  • TextPromptItem – Carries GenerateResponseRequest objects. Produced by TranscriptionNotifier and consumed by LLM handlers including ResponsesApiModelHandler, ChatCompletionsApiModelHandler, and LanguageModelHandler.

  • LMOutItem – Carries LLMResponseChunk, TokenUsage, or EndOfResponse objects. Produced by LLM handlers and consumed by LMOutputProcessor.

  • TTSInItem – Carries TTSInput or EndOfResponse objects. Produced by LMOutputProcessor and consumed by TTS handlers like ChatTTSHandler and Qwen3TTSHandler.

  • AudioOutItem – Carries bytes, np.ndarray, or AudioOutput objects. Produced by TTS handlers and consumed by output sinks like LocalAudioStreamer, WebSocketStreamer, and SocketSender.

  • TextEventItem – Carries PipelineEvent or sentinel bytes. Produced by LMOutputProcessor as a side-channel for real-time text updates and consumed by WebSocket or realtime clients.

This strongly typed system ensures that a mismatch between connected handlers would be caught by static type checkers before runtime.

How Handlers Are Wired in s2s_pipeline.py

The actual connection logic resides in _build_pipeline_handlers inside src/speech_to_speech/s2s_pipeline.py. Each handler receives specific queue instances created by initialize_queues_and_events (lines 34–48).


# From s2s_pipeline.py – handler instantiation showing queue wiring

vad = VADHandler(
    stop_event,
    queue_in=recv_audio_chunks_queue,      # Queue[AudioInItem]

    queue_out=spoken_prompt_queue,         # Queue[VADOutItem]

)

transcription_notifier = TranscriptionNotifier(
    stop_event,
    queue_in=stt_output_queue,             # Queue[STTOutItem]

    queue_out=text_prompt_queue,           # Queue[TextPromptItem]

)

# STT handler receives VADOutItem and emits STTOutItem

stt = get_stt_handler(..., spoken_prompt_queue, stt_output_queue, ...)

# LLM handler consumes TextPromptItem and produces LMOutItem

lm = get_llm_handler(..., text_prompt_queue, lm_response_queue, ...)

# LMOutputProcessor bridges LMOutItem to TTSInItem

lm_processor = LMOutputProcessor(
    stop_event,
    queue_in=lm_response_queue,            # Queue[LMOutItem]

    queue_out=lm_processed_queue,          # Queue[TTSInItem]

    setup_kwargs={"text_output_queue": text_output_queue},  # Queue[TextEventItem]

)

# TTS handler consumes TTSInItem and pushes AudioOutItem

tts = get_tts_handler(..., lm_processed_queue, send_audio_chunks_queue, ...)

This wiring creates a deterministic data flow: AudioInItem → VADOutItem → STTOutItem → TextPromptItem → LMOutItem → TTSInItem → AudioOutItem, with TextEventItem serving as a side channel for real-time text updates.

Practical Pipeline Construction Example

Below is a runnable example demonstrating how to instantiate the typed queues and wire handlers manually. This mirrors the internal logic in s2s_pipeline.py but isolates the components for clarity:

from queue import Queue
from threading import Event
from speech_to_speech.pipeline.queue_types import (
    AudioInItem, VADOutItem, STTOutItem,
    TextPromptItem, LMOutItem, TTSInItem,
    AudioOutItem, TextEventItem,
)
from speech_to_speech.VAD.vad_handler import VADHandler
from speech_to_speech.STT.whisper_stt_handler import WhisperSTTHandler
from speech_to_speech.LLM.language_model import LanguageModelHandler
from speech_to_speech.LLM.lm_output_processor import LMOutputProcessor
from speech_to_speech.TTS.qwen3_tts_handler import Qwen3TTSHandler
from speech_to_speech.utils.thread_manager import ThreadManager

# Initialize typed queues

queues = {
    "audio_in": Queue[AudioInItem](),
    "vad_out": Queue[VADOutItem](),
    "stt_out": Queue[STTOutItem](),
    "text_prompt": Queue[TextPromptItem](),
    "lm_out": Queue[LMOutItem](),
    "tts_in": Queue[TTSInItem](),
    "audio_out": Queue[AudioOutItem](),
    "text_events": Queue[TextEventItem](),
}

stop = Event()
listen = Event()

# Wire handlers with correct queue types

vad = VADHandler(stop, queue_in=queues["audio_in"], queue_out=queues["vad_out"], setup_args=())
stt = WhisperSTTHandler(stop, queue_in=queues["vad_out"], queue_out=queues["stt_out"], setup_kwargs={})
lm = LanguageModelHandler(stop, queue_in=queues["text_prompt"], queue_out=queues["lm_out"], setup_kwargs={})
lm_proc = LMOutputProcessor(
    stop, 
    queue_in=queues["lm_out"], 
    queue_out=queues["tts_in"], 
    setup_kwargs={"text_output_queue": queues["text_events"]}
)
tts = Qwen3TTSHandler(stop, queue_in=queues["tts_in"], queue_out=queues["audio_out"], setup_args=(listen,), setup_kwargs={})

# Execute pipeline

manager = ThreadManager([vad, stt, lm, lm_proc, tts])
manager.start()

# Feed raw audio into queues["audio_in"]...

# manager.stop() on completion

Each handler receives a Queue[<Item>] matching its declared type annotations, ensuring that VADHandler only processes AudioInItem payloads and produces VADOutItem objects for downstream consumers.

Key Source Files

Understanding the queue system requires familiarity with these specific files:

Summary

  • Eight typed queues form the pipeline's backbone: AudioInItem, VADOutItem, STTOutItem, TextPromptItem, LMOutItem, TTSInItem, AudioOutItem, and TextEventItem.
  • Type safety is enforced through TypeAlias definitions in queue_types.py, ensuring handlers receive exactly the data structure they expect.
  • Handler wiring occurs in s2s_pipeline.py via _build_pipeline_handlers, where each component receives Queue[<InItem>] and Queue[<OutItem>] instances.
  • Data flow follows a strict sequence: raw audio enters via AudioInItem, flows through VAD/STT/LLM processing, and exits as AudioOutItem after TTS synthesis.
  • Extensibility is simplified because new handlers only need to declare compatible input/output queue types to integrate with existing pipelines.

Frequently Asked Questions

What is the difference between AudioInItem and AudioOutItem?

AudioInItem carries VADIn payloads (raw audio bytes) entering the pipeline from sources like SocketReceiver or LocalAudioStreamer, and is consumed exclusively by VADHandler. AudioOutItem carries processed audio data (bytes, np.ndarray, or AudioOutput objects) leaving the TTS stage, consumed by output sinks like SocketSender or playback streams.

How does LMOutputProcessor bridge the LLM and TTS stages?

LMOutputProcessor (in src/speech_to_speech/LLM/lm_output_processor.py) consumes LMOutItem from the LLM handler, which may contain LLMResponseChunk, TokenUsage, or EndOfResponse objects. It processes these into TTSInput instances wrapped as TTSInItem, pushing them to the TTS queue. Additionally, it emits TextEventItem objects to a side-channel queue for real-time text streaming to clients.

Where are the queue types defined in the source code?

All queue type aliases are defined in src/speech_to_speech/pipeline/queue_types.py as TypeAlias assignments. The concrete message classes that populate these queues (like VADAudio or GenerateResponseRequest) are defined in src/speech_to_speech/pipeline/messages.py.

Why does the pipeline use typed queues instead of raw byte streams?

Typed queues provide compile-time safety and explicit data contracts between stages. Because each handler declares Queue[AudioInItem] or Queue[TTSInItem] in its constructor, static type checkers can verify that VADHandler never receives TTS output accidentally, and developers can trace the exact data shape flowing through the system by examining the TypeAlias definitions.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →