# Speech-to-Speech vs Text-to-Speech: Understanding the Key Difference in the Hugging Face Pipeline

> Learn the key difference between speech-to-speech and text-to-speech in the Hugging Face pipeline. Speech-to-speech is a full voice pipeline, while text-to-speech is just the final synthesis step.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: deep-dive
- Published: 2026-08-01

---

**Speech-to-speech is a complete voice-agent pipeline that chains four components—Voice Activity Detection, Speech-to-Text, Language Model, and Text-to-Speech—to turn spoken input into spoken responses, while text-to-speech is only the final synthesis stage that converts text into audio.**

Understanding the difference between speech-to-speech and text-to-speech is essential when working with the `huggingface/speech-to-speech` repository. While both technologies produce audio output, they differ fundamentally in scope, input types, and processing complexity. This guide breaks down the architectural distinctions using the actual source code implementation to show exactly where each technology fits in the voice AI stack.

## What Is Speech-to-Speech?

Speech-to-speech (S2S) refers to the complete **end-to-end voice agent system** implemented in [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py). It is a full-stack pipeline that listens to spoken user utterances, processes them through multiple AI stages, and returns synthesized speech responses.

### The Four-Stage Pipeline

The speech-to-speech system orchestrates four interchangeable handlers, each running in its own thread and connected by message queues:

1. **Voice Activity Detection (VAD)** – Implemented in [`src/speech_to_speech/VAD/vad_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/VAD/vad_handler.py), this component uses Silero VAD to detect when users start and stop speaking, enabling real-time turn-taking.
2. **Speech-to-Text (STT)** – The `ParakeetTDTHandler` class in [`src/speech_to_speech/STT/parakeet_tdt_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/STT/parakeet_tdt_handler.py) transcribes raw audio into text.
3. **Language Model (LLM)** – Located in [`src/speech_to_speech/LLM/chat.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/chat.py), this stage generates textual responses and handles tool calls.
4. **Text-to-Speech (TTS)** – The final synthesis stage that converts the LLM's text output back into audio.

Because S2S integrates all four stages, it can handle **partial transcripts**, **conversational context**, and **tool calling** that pure audio synthesis cannot manage.

## What Is Text-to-Speech?

Text-to-speech (TTS) is **only the final audio generation stage** of the larger pipeline. In the `huggingface/speech-to-speech` repository, TTS implementations live in `src/speech_to_speech/TTS/` and receive plain text (or pre-processed prompts) as input.

The primary implementation, `Qwen3TTSHandler` in [`src/speech_to_speech/TTS/qwen3_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/qwen3_tts_handler.py), supports two distinct backends:
- **mlx-audio** for Apple Silicon devices
- **faster-qwen3-tts** for CUDA and CPU platforms

Unlike the full speech-to-speech system, a standalone TTS handler cannot interpret spoken input or generate conversational responses—it merely synthesizes whatever text it receives.

## Key Differences Between Speech-to-Speech and Text-to-Speech

The distinction comes down to **architectural scope**:

| Aspect | Speech-to-Speech | Text-to-Speech |
|--------|------------------|----------------|
| **Input** | Audio (speech) | Text |
| **Processing** | VAD → STT → LLM → TTS | Only TTS conversion |
| **Output** | Audio (speech) | Audio (speech) |
| **Capabilities** | Turn-taking, context awareness, tool use | Voice synthesis only |
| **Primary File** | [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py) | [`src/speech_to_speech/TTS/qwen3_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/qwen3_tts_handler.py) |

Because speech-to-speech incorporates the STT and LLM stages, it functions as a **conversational agent**, while text-to-speech serves as a **narration utility** for pre-generated content.

## Practical Code Examples

### Running the Full Speech-to-Speech Pipeline

To launch the complete four-stage pipeline with default components:

```bash

# Install the package

pip install speech-to-speech

# Start the realtime-compatible server (VAD → STT → LLM → TTS)

speech-to-speech

```

This command initializes the orchestration logic in [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py), which wires all handlers together via the message classes defined in [`src/speech_to_speech/pipeline/messages.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/pipeline/messages.py).

### Using the TTS Handler Standalone

You can import and use only the TTS component without the rest of the pipeline:

```python
from speech_to_speech.TTS.qwen3_tts_handler import Qwen3TTSHandler
from speech_to_speech.pipeline.handler_types import TTSIn, TTSOut
from threading import Event

# Create a dummy event (required by the BaseHandler)

listen_event = Event()
listen_event.set()

# Initialise the handler (defaults load the Qwen3-TTS model)

handler = Qwen3TTSHandler()
handler.setup(should_listen=listen_event)

# Prepare a simple TTS input

tts_input = TTSIn(text="Hello, world!", language="en", speaker="Aiden")

# Run the synthesis (returns an iterator of audio chunks)

audio_chunks = handler.handle(tts_input)

# Consume the chunks (e.g., write to a WAV file)

with open("hello.wav", "wb") as f:
    for chunk in audio_chunks:
        f.write(chunk.audio)

```

This example demonstrates how TTS operates in isolation, receiving text through the `TTSIn` class and returning audio chunks without any speech recognition or language understanding.

### Using the STT Component Independently

Similarly, you can use just the speech recognition stage:

```python
from speech_to_speech.STT.parakeet_tdt_handler import ParakeetTDTHandler
from speech_to_speech.pipeline.handler_types import STTIn, STTOut
from threading import Event

listen_event = Event()
listen_event.set()

stt_handler = ParakeetTDTHandler()
stt_handler.setup(should_listen=listen_event)

audio_input = STTIn(audio_path="speech.wav")
transcript = stt_handler.handle(audio_input)  # returns STTOut with .text

print(transcript.text)

```

## Summary

- **Speech-to-speech** is an end-to-end pipeline ([`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py)) that combines VAD, STT, LLM, and TTS to create conversational voice agents.
- **Text-to-speech** is only the final synthesis stage ([`src/speech_to_speech/TTS/qwen3_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/qwen3_tts_handler.py)) that converts text to audio.
- The key difference is architectural scope: S2S processes audio input through multiple AI stages, while TTS requires text input and performs only voice generation.
- All components are modular, allowing developers to use the full pipeline or import individual handlers (STT, TTS) for specific use cases.

## Frequently Asked Questions

### Can I use the TTS module without the full speech-to-speech pipeline?

Yes. The `Qwen3TTSHandler` class in [`src/speech_to_speech/TTS/qwen3_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/qwen3_tts_handler.py) can be instantiated and used independently, as shown in the standalone example above. You only need to provide a `TTSIn` object containing the text, language, and speaker parameters.

### What hardware backends does the TTS handler support?

According to the source code in [`qwen3_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/qwen3_tts_handler.py) (lines 80-88), the handler automatically selects between **mlx-audio** for Apple Silicon devices and **faster-qwen3-tts** for CUDA and CPU platforms. This detection happens during the `setup()` method call.

### How does the VAD component enable real-time conversation?

The `VADHandler` in [`src/speech_to_speech/VAD/vad_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/VAD/vad_handler.py) uses Silero VAD to detect speech boundaries. It feeds turn-taking information to the pipeline, allowing the system to distinguish between user speech and system responses, which is essential for natural dialogue that pure TTS cannot achieve.

### Which handler orchestrates the pipeline components?

The [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py) file contains the main orchestration logic that instantiates all four handlers (VAD, STT, LLM, TTS), manages their threading, and connects them via queues using message types defined in [`src/speech_to_speech/pipeline/messages.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/pipeline/messages.py).