Speech-to-Speech vs Text-to-Speech: Understanding the Key Difference in the Hugging Face Pipeline

Speech-to-speech is a complete voice-agent pipeline that chains four components—Voice Activity Detection, Speech-to-Text, Language Model, and Text-to-Speech—to turn spoken input into spoken responses, while text-to-speech is only the final synthesis stage that converts text into audio.

Understanding the difference between speech-to-speech and text-to-speech is essential when working with the huggingface/speech-to-speech repository. While both technologies produce audio output, they differ fundamentally in scope, input types, and processing complexity. This guide breaks down the architectural distinctions using the actual source code implementation to show exactly where each technology fits in the voice AI stack.

What Is Speech-to-Speech?

Speech-to-speech (S2S) refers to the complete end-to-end voice agent system implemented in src/speech_to_speech/s2s_pipeline.py. It is a full-stack pipeline that listens to spoken user utterances, processes them through multiple AI stages, and returns synthesized speech responses.

The Four-Stage Pipeline

The speech-to-speech system orchestrates four interchangeable handlers, each running in its own thread and connected by message queues:

  1. Voice Activity Detection (VAD) – Implemented in src/speech_to_speech/VAD/vad_handler.py, this component uses Silero VAD to detect when users start and stop speaking, enabling real-time turn-taking.
  2. Speech-to-Text (STT) – The ParakeetTDTHandler class in src/speech_to_speech/STT/parakeet_tdt_handler.py transcribes raw audio into text.
  3. Language Model (LLM) – Located in src/speech_to_speech/LLM/chat.py, this stage generates textual responses and handles tool calls.
  4. Text-to-Speech (TTS) – The final synthesis stage that converts the LLM's text output back into audio.

Because S2S integrates all four stages, it can handle partial transcripts, conversational context, and tool calling that pure audio synthesis cannot manage.

What Is Text-to-Speech?

Text-to-speech (TTS) is only the final audio generation stage of the larger pipeline. In the huggingface/speech-to-speech repository, TTS implementations live in src/speech_to_speech/TTS/ and receive plain text (or pre-processed prompts) as input.

The primary implementation, Qwen3TTSHandler in src/speech_to_speech/TTS/qwen3_tts_handler.py, supports two distinct backends:

  • mlx-audio for Apple Silicon devices
  • faster-qwen3-tts for CUDA and CPU platforms

Unlike the full speech-to-speech system, a standalone TTS handler cannot interpret spoken input or generate conversational responses—it merely synthesizes whatever text it receives.

Key Differences Between Speech-to-Speech and Text-to-Speech

The distinction comes down to architectural scope:

Aspect Speech-to-Speech Text-to-Speech
Input Audio (speech) Text
Processing VAD → STT → LLM → TTS Only TTS conversion
Output Audio (speech) Audio (speech)
Capabilities Turn-taking, context awareness, tool use Voice synthesis only
Primary File src/speech_to_speech/s2s_pipeline.py src/speech_to_speech/TTS/qwen3_tts_handler.py

Because speech-to-speech incorporates the STT and LLM stages, it functions as a conversational agent, while text-to-speech serves as a narration utility for pre-generated content.

Practical Code Examples

Running the Full Speech-to-Speech Pipeline

To launch the complete four-stage pipeline with default components:


# Install the package

pip install speech-to-speech

# Start the realtime-compatible server (VAD → STT → LLM → TTS)

speech-to-speech

This command initializes the orchestration logic in src/speech_to_speech/s2s_pipeline.py, which wires all handlers together via the message classes defined in src/speech_to_speech/pipeline/messages.py.

Using the TTS Handler Standalone

You can import and use only the TTS component without the rest of the pipeline:

from speech_to_speech.TTS.qwen3_tts_handler import Qwen3TTSHandler
from speech_to_speech.pipeline.handler_types import TTSIn, TTSOut
from threading import Event

# Create a dummy event (required by the BaseHandler)

listen_event = Event()
listen_event.set()

# Initialise the handler (defaults load the Qwen3-TTS model)

handler = Qwen3TTSHandler()
handler.setup(should_listen=listen_event)

# Prepare a simple TTS input

tts_input = TTSIn(text="Hello, world!", language="en", speaker="Aiden")

# Run the synthesis (returns an iterator of audio chunks)

audio_chunks = handler.handle(tts_input)

# Consume the chunks (e.g., write to a WAV file)

with open("hello.wav", "wb") as f:
    for chunk in audio_chunks:
        f.write(chunk.audio)

This example demonstrates how TTS operates in isolation, receiving text through the TTSIn class and returning audio chunks without any speech recognition or language understanding.

Using the STT Component Independently

Similarly, you can use just the speech recognition stage:

from speech_to_speech.STT.parakeet_tdt_handler import ParakeetTDTHandler
from speech_to_speech.pipeline.handler_types import STTIn, STTOut
from threading import Event

listen_event = Event()
listen_event.set()

stt_handler = ParakeetTDTHandler()
stt_handler.setup(should_listen=listen_event)

audio_input = STTIn(audio_path="speech.wav")
transcript = stt_handler.handle(audio_input)  # returns STTOut with .text

print(transcript.text)

Summary

  • Speech-to-speech is an end-to-end pipeline (src/speech_to_speech/s2s_pipeline.py) that combines VAD, STT, LLM, and TTS to create conversational voice agents.
  • Text-to-speech is only the final synthesis stage (src/speech_to_speech/TTS/qwen3_tts_handler.py) that converts text to audio.
  • The key difference is architectural scope: S2S processes audio input through multiple AI stages, while TTS requires text input and performs only voice generation.
  • All components are modular, allowing developers to use the full pipeline or import individual handlers (STT, TTS) for specific use cases.

Frequently Asked Questions

Can I use the TTS module without the full speech-to-speech pipeline?

Yes. The Qwen3TTSHandler class in src/speech_to_speech/TTS/qwen3_tts_handler.py can be instantiated and used independently, as shown in the standalone example above. You only need to provide a TTSIn object containing the text, language, and speaker parameters.

What hardware backends does the TTS handler support?

According to the source code in qwen3_tts_handler.py (lines 80-88), the handler automatically selects between mlx-audio for Apple Silicon devices and faster-qwen3-tts for CUDA and CPU platforms. This detection happens during the setup() method call.

How does the VAD component enable real-time conversation?

The VADHandler in src/speech_to_speech/VAD/vad_handler.py uses Silero VAD to detect speech boundaries. It feeds turn-taking information to the pipeline, allowing the system to distinguish between user speech and system responses, which is essential for natural dialogue that pure TTS cannot achieve.

Which handler orchestrates the pipeline components?

The s2s_pipeline.py file contains the main orchestration logic that instantiates all four handlers (VAD, STT, LLM, TTS), manages their threading, and connects them via queues using message types defined in src/speech_to_speech/pipeline/messages.py.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →