# How Does the Speech-to-Speech Model Generate Output? A Deep Dive into the Pipeline

> Explore the speech-to-speech model pipeline: VAD, STT, LLM, processing, TTS, and streaming. Understand how audio input becomes spoken output.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: deep-dive
- Published: 2026-07-07

---

**The speech-to-speech model generates output through a modular, asynchronous pipeline that converts audio input into spoken responses via six sequential stages: Voice Activity Detection (VAD), Speech-to-Text (STT), Language Model (LLM) inference, output processing, Text-to-Speech (TTS), and audio streaming.**

The `huggingface/speech-to-speech` repository implements this architecture as a series of interconnected handlers that communicate through thread-safe queues. Understanding how the speech-to-speech model generates output requires examining the data flow from raw audio capture to final PCM audio playback.

## The Six-Stage Generation Pipeline

### Stage 1: Voice Activity Detection (VAD)

The pipeline begins with the `VADHandler` in [`speech_to_speech/VAD/vad_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/speech_to_speech/VAD/vad_handler.py), which processes incoming audio streams from microphones, sockets, or local files. This handler trims silence and emits spoken prompts—chunks of voice activity—into the pipeline for downstream processing.

### Stage 2: Speech-to-Text Transcription (STT)

Each spoken prompt flows to a selected STT backend following the `BaseSTTHandler` contract. The concrete implementations include Whisper, Faster-Whisper, Paraformer, and others located in [`speech_to_speech/STT/whisper_stt_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/speech_to_speech/STT/whisper_stt_handler.py). These handlers return `LLMResponseChunk` objects containing transcribed text and optional tool-call metadata.

### Stage 3: Language Model Processing

The transcribed text passes to one of three language model handlers: `LanguageModelHandler`, `ResponsesApiModelHandler`, or `ChatCompletionsApiModelHandler` in [`speech_to_speech/LLM/language_model.py`](https://github.com/huggingface/speech-to-speech/blob/main/speech_to_speech/LLM/language_model.py). The LLM streams back `LLMResponseChunk` objects, `TokenUsage` reports, and a final `EndOfResponse` signal to indicate completion.

### Stage 4: LM Output Processing

The `LMOutputProcessor` in [`speech_to_speech/LLM/lm_output_processor.py`](https://github.com/huggingface/speech-to-speech/blob/main/speech_to_speech/LLM/lm_output_processor.py) intercepts the LLM stream and performs two critical functions:

- Pushes `AssistantTextEvent` objects (text + optional tool calls) onto a text-output queue for WebSocket clients
- Yields clean `TTSInput` objects downstream only when `response_wants_audio` is true

The processor uses `SpeculativeTurnTracker` to drop stale turns and prevent "ghost" generations from outdated responses.

### Stage 5: Text-to-Speech Synthesis

The `TTSInput` flows to concrete TTS handlers such as `Qwen3TTSHandler`, `ChatTTSHandler`, or `PocketTTSHandler`. The implementation in [`speech_to_speech/TTS/qwen3_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/speech_to_speech/TTS/qwen3_tts_handler.py) streams int16 PCM audio chunks while handling voice-cloning, custom-speaker, or voice-design modes. These handlers respect session-level voice overrides from OpenAI Realtime responses or runtime configuration.

### Stage 6: Audio Output Streaming

Generated audio chunks enter the send-audio queue, consumed by either `LocalAudioStreamer` or `SocketSender`. The [`speech_to_speech/connections/local_audio_streamer.py`](https://github.com/huggingface/speech-to-speech/blob/main/speech_to_speech/connections/local_audio_streamer.py) implementation plays audio on local sound devices or transmits it over the network, completing the generation cycle.

## Pipeline Orchestration and Queue Management

### Queue Types and Thread Safety

All components communicate via thread-safe `Queue` objects defined in [`speech_to_speech/pipeline/queue_types.py`](https://github.com/huggingface/speech-to-speech/blob/main/speech_to_speech/pipeline/queue_types.py). The `initialize_queues_and_events` helper creates queues for:

- Stop signals
- VAD output
- STT output
- Text prompts
- LLM responses
- TTS input
- Final audio

### Speculative Turn Tracking

For real-time usage, the `SpeculativeTurnTracker` in [`speech_to_speech/pipeline/speculative_turns.py`](https://github.com/huggingface/speech-to-speech/blob/main/speech_to_speech/pipeline/speculative_turns.py) tracks turn IDs to silently discard late-arriving STT or LLM chunks belonging to older turns. This prevents the system from speaking outdated content when users interrupt or when network latency causes out-of-order processing.

### Runtime Configuration

In OpenAI Realtime mode, the system accepts a `RealtimeSessionCreateRequest` supplied to LLM and TTS handlers via [`speech_to_speech/api/openai_realtime/runtime_config.py`](https://github.com/huggingface/speech-to-speech/blob/main/speech_to_speech/api/openai_realtime/runtime_config.py). This enables session-level voice overrides and streaming controls without restarting the pipeline.

## Running the Pipeline

### Command-Line Execution

Execute the full pipeline using the [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py) module:

```bash
python -m speech_to_speech.s2s_pipeline \
  --mode realtime \
  --stt whisper \
  --llm_backend responses-api \
  --tts qwen3 \
  --ws_host localhost \
  --ws_port 8080

```

### Programmatic Python API

Build and start the pipeline programmatically:

```python
from speech_to_speech.s2s_pipeline import (
    parse_arguments,
    prepare_all_args,
    initialize_queues_and_events,
    build_pipeline,
)

# Parse CLI-style arguments or construct manually

args = parse_arguments()
prepare_all_args(
    args.module_kwargs,
    args.whisper_stt_handler_kwargs,
    args.language_model_handler_kwargs,
    args.qwen3_tts_handler_kwargs,
    # ... other handler kwargs

)

# Build queues and pipeline

queues = initialize_queues_and_events()
pipeline = build_pipeline(
    args.module_kwargs,
    args.socket_receiver_kwargs,
    args.socket_sender_kwargs,
    args.vad_handler_kwargs,
    args.whisper_stt_handler_kwargs,
    args.language_model_handler_kwargs,
    args.qwen3_tts_handler_kwargs,
    queues,
)

pipeline.start()  # Launches all background threads

pipeline.wait()   # Blocks until shutdown

```

### Local Audio Testing

For local testing without network configuration:

```bash
python scripts/listen_and_play.py \
  --send_rate 16000 \
  --recv_rate 16000 \
  --host localhost \
  --send_port 12345 \
  --recv_port 12346

```

This opens socket pairs, streams microphone audio to the pipeline, and plays generated speech in real-time.

## Summary

- The speech-to-speech model generates output through a **six-stage asynchronous pipeline** (VAD → STT → LLM → Processor → TTS → Audio).
- **Queue-based architecture** in [`queue_types.py`](https://github.com/huggingface/speech-to-speech/blob/main/queue_types.py) ensures thread-safe communication between handlers without blocking.
- **Speculative turn tracking** prevents outdated responses from being spoken when users interrupt or network latency occurs.
- **Modular handler design** allows swapping STT backends (Whisper, Paraformer), LLM providers, and TTS engines (Qwen3, ChatTTS) without changing core logic.
- **OpenAI Realtime compatibility** enables session-level voice configuration and streaming controls via [`runtime_config.py`](https://github.com/huggingface/speech-to-speech/blob/main/runtime_config.py).

## Frequently Asked Questions

### What prevents the speech-to-speech model from playing outdated responses?

The `SpeculativeTurnTracker` in [`speech_to_speech/pipeline/speculative_turns.py`](https://github.com/huggingface/speech-to-speech/blob/main/speech_to_speech/pipeline/speculative_turns.py) tracks unique turn IDs across the pipeline. When new user input arrives, older turns are marked stale, and the `LMOutputProcessor` silently drops any late-arriving chunks from previous turns before they reach the TTS stage.

### How does the pipeline support different TTS backends?

The architecture uses a base handler pattern where concrete implementations like `Qwen3TTSHandler` and `ChatTTSHandler` inherit from the base class defined in [`speech_to_speech/baseHandler.py`](https://github.com/huggingface/speech-to-speech/blob/main/speech_to_speech/baseHandler.py). Each handler receives `TTSInput` objects and streams int16 PCM audio, allowing the `build_pipeline` function in [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py) to inject the appropriate backend based on runtime arguments.

### Can the speech-to-speech model operate without speaking responses?

Yes. The `LMOutputProcessor` checks the `response_wants_audio` flag before yielding `TTSInput` downstream. If disabled, the system still pushes `AssistantTextEvent` objects to the text-output queue for WebSocket clients, enabling text-only responses while maintaining the full LLM inference capability.

### What audio format does the LocalAudioStreamer expect?

According to [`speech_to_speech/connections/local_audio_streamer.py`](https://github.com/huggingface/speech-to-speech/blob/main/speech_to_speech/connections/local_audio_streamer.py), the system expects **int16 PCM audio chunks** at the sample rate specified during initialization (typically 16kHz or 24kHz). The TTS handlers convert their native outputs to this format before placing chunks on the send-audio queue.