# Speech-to-Speech Queue Types: How AudioInItem, TTSInItem, and Typed Queues Connect Pipeline Handlers

> Understand speech-to-speech queue types like AudioInItem and TTSInItem. Learn how these typed queues ensure type-safe data flow between pipeline handlers in the Hugging Face library.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: internals
- Published: 2026-07-11

---

**The speech-to-speech library uses eight strongly-typed queue aliases defined in [`queue_types.py`](https://github.com/huggingface/speech-to-speech/blob/main/queue_types.py)—including `AudioInItem`, `VADOutItem`, `STTOutItem`, `TextPromptItem`, `LMOutItem`, `TTSInItem`, `AudioOutItem`, and `TextEventItem`—to enforce type-safe data flow between handlers, with each stage receiving a `Queue[<InItem>]` and emitting to a `Queue[<OutItem>]`.**

The Hugging Face `speech-to-speech` repository implements a modular real-time voice processing pipeline where data moves through consecutive stages via Python's `Queue` objects. Unlike generic byte streams, the system employs **typed queues** that carry specific payload objects, making the data flow explicit and statically verifiable. This architecture centers on [`src/speech_to_speech/pipeline/queue_types.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/pipeline/queue_types.py), which defines type aliases that determine exactly what data travels between the VAD, STT, LLM, and TTS handlers.

## The Eight Queue Types Defined in queue_types.py

The pipeline's type safety relies on eight distinct `TypeAlias` definitions in [`src/speech_to_speech/pipeline/queue_types.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/pipeline/queue_types.py). Each alias represents the exact payload shape a handler expects on its input and produces on its output.

- **`AudioInItem`** – Carries `VADIn` payloads (raw audio bytes) entering from `SocketReceiver`, `WebSocketStreamer`, or `LocalAudioStreamer`. Consumed exclusively by `VADHandler`.

- **`VADOutItem`** – Carries `VADAudio` objects (voice activity segments). Produced by `VADHandler` and consumed by STT handlers like `WhisperSTTHandler`.

- **`STTOutItem`** – Carries `PartialTranscription` or `Transcription` objects. Produced by any STT handler and consumed by `TranscriptionNotifier`.

- **`TextPromptItem`** – Carries `GenerateResponseRequest` objects. Produced by `TranscriptionNotifier` and consumed by LLM handlers including `ResponsesApiModelHandler`, `ChatCompletionsApiModelHandler`, and `LanguageModelHandler`.

- **`LMOutItem`** – Carries `LLMResponseChunk`, `TokenUsage`, or `EndOfResponse` objects. Produced by LLM handlers and consumed by `LMOutputProcessor`.

- **`TTSInItem`** – Carries `TTSInput` or `EndOfResponse` objects. Produced by `LMOutputProcessor` and consumed by TTS handlers like `ChatTTSHandler` and `Qwen3TTSHandler`.

- **`AudioOutItem`** – Carries `bytes`, `np.ndarray`, or `AudioOutput` objects. Produced by TTS handlers and consumed by output sinks like `LocalAudioStreamer`, `WebSocketStreamer`, and `SocketSender`.

- **`TextEventItem`** – Carries `PipelineEvent` or sentinel `bytes`. Produced by `LMOutputProcessor` as a side-channel for real-time text updates and consumed by WebSocket or realtime clients.

This strongly typed system ensures that a mismatch between connected handlers would be caught by static type checkers before runtime.

## How Handlers Are Wired in s2s_pipeline.py

The actual connection logic resides in `_build_pipeline_handlers` inside [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py). Each handler receives specific queue instances created by `initialize_queues_and_events` (lines 34–48).

```python

# From s2s_pipeline.py – handler instantiation showing queue wiring

vad = VADHandler(
    stop_event,
    queue_in=recv_audio_chunks_queue,      # Queue[AudioInItem]

    queue_out=spoken_prompt_queue,         # Queue[VADOutItem]

)

transcription_notifier = TranscriptionNotifier(
    stop_event,
    queue_in=stt_output_queue,             # Queue[STTOutItem]

    queue_out=text_prompt_queue,           # Queue[TextPromptItem]

)

# STT handler receives VADOutItem and emits STTOutItem

stt = get_stt_handler(..., spoken_prompt_queue, stt_output_queue, ...)

# LLM handler consumes TextPromptItem and produces LMOutItem

lm = get_llm_handler(..., text_prompt_queue, lm_response_queue, ...)

# LMOutputProcessor bridges LMOutItem to TTSInItem

lm_processor = LMOutputProcessor(
    stop_event,
    queue_in=lm_response_queue,            # Queue[LMOutItem]

    queue_out=lm_processed_queue,          # Queue[TTSInItem]

    setup_kwargs={"text_output_queue": text_output_queue},  # Queue[TextEventItem]

)

# TTS handler consumes TTSInItem and pushes AudioOutItem

tts = get_tts_handler(..., lm_processed_queue, send_audio_chunks_queue, ...)

```

This wiring creates a deterministic data flow: **AudioInItem → VADOutItem → STTOutItem → TextPromptItem → LMOutItem → TTSInItem → AudioOutItem**, with `TextEventItem` serving as a side channel for real-time text updates.

## Practical Pipeline Construction Example

Below is a runnable example demonstrating how to instantiate the typed queues and wire handlers manually. This mirrors the internal logic in [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py) but isolates the components for clarity:

```python
from queue import Queue
from threading import Event
from speech_to_speech.pipeline.queue_types import (
    AudioInItem, VADOutItem, STTOutItem,
    TextPromptItem, LMOutItem, TTSInItem,
    AudioOutItem, TextEventItem,
)
from speech_to_speech.VAD.vad_handler import VADHandler
from speech_to_speech.STT.whisper_stt_handler import WhisperSTTHandler
from speech_to_speech.LLM.language_model import LanguageModelHandler
from speech_to_speech.LLM.lm_output_processor import LMOutputProcessor
from speech_to_speech.TTS.qwen3_tts_handler import Qwen3TTSHandler
from speech_to_speech.utils.thread_manager import ThreadManager

# Initialize typed queues

queues = {
    "audio_in": Queue[AudioInItem](),
    "vad_out": Queue[VADOutItem](),
    "stt_out": Queue[STTOutItem](),
    "text_prompt": Queue[TextPromptItem](),
    "lm_out": Queue[LMOutItem](),
    "tts_in": Queue[TTSInItem](),
    "audio_out": Queue[AudioOutItem](),
    "text_events": Queue[TextEventItem](),
}

stop = Event()
listen = Event()

# Wire handlers with correct queue types

vad = VADHandler(stop, queue_in=queues["audio_in"], queue_out=queues["vad_out"], setup_args=())
stt = WhisperSTTHandler(stop, queue_in=queues["vad_out"], queue_out=queues["stt_out"], setup_kwargs={})
lm = LanguageModelHandler(stop, queue_in=queues["text_prompt"], queue_out=queues["lm_out"], setup_kwargs={})
lm_proc = LMOutputProcessor(
    stop, 
    queue_in=queues["lm_out"], 
    queue_out=queues["tts_in"], 
    setup_kwargs={"text_output_queue": queues["text_events"]}
)
tts = Qwen3TTSHandler(stop, queue_in=queues["tts_in"], queue_out=queues["audio_out"], setup_args=(listen,), setup_kwargs={})

# Execute pipeline

manager = ThreadManager([vad, stt, lm, lm_proc, tts])
manager.start()

# Feed raw audio into queues["audio_in"]...

# manager.stop() on completion

```

Each handler receives a `Queue[<Item>]` matching its declared type annotations, ensuring that `VADHandler` only processes `AudioInItem` payloads and produces `VADOutItem` objects for downstream consumers.

## Key Source Files

Understanding the queue system requires familiarity with these specific files:

- **[`src/speech_to_speech/pipeline/queue_types.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/pipeline/queue_types.py)** – Central definition of all queue payload aliases (`AudioInItem`, `VADOutItem`, etc.)
- **[`src/speech_to_speech/pipeline/handler_types.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/pipeline/handler_types.py)** – Generic handler signatures and protocol definitions
- **[`src/speech_to_speech/pipeline/messages.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/pipeline/messages.py)** – Concrete Pydantic message classes (e.g., `VADAudio`, `Transcription`, `TTSInput`)
- **[`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py)** – Orchestrates queue creation (`initialize_queues_and_events`) and handler wiring (`_build_pipeline_handlers`)
- **[`src/speech_to_speech/LLM/lm_output_processor.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/lm_output_processor.py)** – Bridges `LMOutItem` to `TTSInItem` and emits `TextEventItem` side channels

## Summary

- **Eight typed queues** form the pipeline's backbone: `AudioInItem`, `VADOutItem`, `STTOutItem`, `TextPromptItem`, `LMOutItem`, `TTSInItem`, `AudioOutItem`, and `TextEventItem`.
- **Type safety** is enforced through `TypeAlias` definitions in [`queue_types.py`](https://github.com/huggingface/speech-to-speech/blob/main/queue_types.py), ensuring handlers receive exactly the data structure they expect.
- **Handler wiring** occurs in [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py) via `_build_pipeline_handlers`, where each component receives `Queue[<InItem>]` and `Queue[<OutItem>]` instances.
- **Data flow** follows a strict sequence: raw audio enters via `AudioInItem`, flows through VAD/STT/LLM processing, and exits as `AudioOutItem` after TTS synthesis.
- **Extensibility** is simplified because new handlers only need to declare compatible input/output queue types to integrate with existing pipelines.

## Frequently Asked Questions

### What is the difference between AudioInItem and AudioOutItem?

`AudioInItem` carries `VADIn` payloads (raw audio bytes) entering the pipeline from sources like `SocketReceiver` or `LocalAudioStreamer`, and is consumed exclusively by `VADHandler`. `AudioOutItem` carries processed audio data (`bytes`, `np.ndarray`, or `AudioOutput` objects) leaving the TTS stage, consumed by output sinks like `SocketSender` or playback streams.

### How does LMOutputProcessor bridge the LLM and TTS stages?

`LMOutputProcessor` (in [`src/speech_to_speech/LLM/lm_output_processor.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/lm_output_processor.py)) consumes `LMOutItem` from the LLM handler, which may contain `LLMResponseChunk`, `TokenUsage`, or `EndOfResponse` objects. It processes these into `TTSInput` instances wrapped as `TTSInItem`, pushing them to the TTS queue. Additionally, it emits `TextEventItem` objects to a side-channel queue for real-time text streaming to clients.

### Where are the queue types defined in the source code?

All queue type aliases are defined in [`src/speech_to_speech/pipeline/queue_types.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/pipeline/queue_types.py) as `TypeAlias` assignments. The concrete message classes that populate these queues (like `VADAudio` or `GenerateResponseRequest`) are defined in [`src/speech_to_speech/pipeline/messages.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/pipeline/messages.py).

### Why does the pipeline use typed queues instead of raw byte streams?

Typed queues provide **compile-time safety** and **explicit data contracts** between stages. Because each handler declares `Queue[AudioInItem]` or `Queue[TTSInItem]` in its constructor, static type checkers can verify that `VADHandler` never receives TTS output accidentally, and developers can trace the exact data shape flowing through the system by examining the `TypeAlias` definitions.