How to Run a Basic Speech-to-Speech Example with HuggingFace: Complete Setup Guide
Run a real-time speech-to-speech pipeline using python -m scripts.listen_and_play --stt whisper_tiny --llm openai --tts pocket_tts after installing the package and setting your API keys.
The HuggingFace speech-to-speech repository provides a modular, real-time pipeline that chains together speech recognition (STT), language generation (LLM), and speech synthesis (TTS) into a single streaming system. This guide walks you through running the minimal "listen-and-play" demo that captures your microphone audio, processes it through all three stages, and plays back the synthesized response.
Understanding the Pipeline Architecture
The speech-to-speech system operates as a coordinated chain of specialized handlers orchestrated by S2SPipeline in src/speech_to_speech/s2s_pipeline.py. Understanding this flow helps you troubleshoot and customize your setup.
Data Flow Through the System
- Audio Capture —
local_audio_streamer.pyreads microphone input continuously - Voice Detection —
vad_iterator.pyidentifies speech segments to process - Speech-to-Text — One of
STT/*_handler.pymodules (Whisper, Paraformer, MMS) transcribes audio - Language Generation — An LLM backend (
responses_api_language_model.py,chat_completions_language_model.py) generates responses - Text-to-Speech — A TTS handler (
pocket_tts_handler.py,qwen3_tts_handler.py, etc.) synthesizes audio - Playback —
listen_and_play.pysends audio directly to your speakers or streams viawebsocket_streamer.py
All components share a unified configuration through typed Arguments classes defined in src/speech_to_speech/arguments_classes/*_arguments.py, making backend swaps as simple as changing a CLI flag.
Installation and Prerequisites
Step 1: Install the Package
The repository supports optional dependencies for different model backends. For the simplest setup, install with all extras:
pip install "speech-to-speech[all]"
For a lighter installation, you can specify only the backends you need.
Step 2: Configure API Credentials
Most LLM backends require authentication. Export your keys before running:
export OPENAI_API_KEY=your_openai_key_here
Other backends may need additional environment variables—check the handler's argument class for specifics.
Running Your First Speech-to-Speech Demo
The Minimal Command
Execute the end-to-end pipeline with three CLI flags selecting your backends:
python -m scripts.listen_and_play \
--stt whisper_tiny \
--llm openai \
--tts pocket_tts
What happens when you run this:
scripts/listen_and_play.pyinstantiatesS2SPipelinewithArgumentsparsed from your flags- The pipeline creates handlers:
whisper_handler.pyfor STT, an OpenAI-compatible LLM client, andpocket_tts_handler.pyfor synthesis local_audio_streamer.pybegins capturing microphone audio through the VAD- Each detected speech segment flows through STT → LLM → TTS → local audio output
Speak into your microphone—you'll hear your synthesized response within seconds.
Alternative Backend Combinations
Swap any component by changing its flag. The pipeline automatically loads the corresponding handler from src/speech_to_speech/:
# Use Paraformer for STT, Responses API for LLM, Qwen-3 for TTS
python -m scripts.listen_and_play \
--stt paraformer \
--llm responses_api \
--tts qwen3
Available options are defined in the argument classes and include:
- STT:
whisper_tiny,whisper_base,whisper_large,paraformer,mms - LLM:
openai,responses_api,chat_completions(plus any OpenAI-compatible endpoint) - TTS:
pocket_tts,qwen3,kokoro,chattts
Key Source Files for Reference
| File Path | Purpose |
|---|---|
scripts/listen_and_play.py |
Entry-point demo script that wires all components together |
src/speech_to_speech/s2s_pipeline.py |
Core orchestration; creates handlers and manages the processing loop |
src/speech_to_speech/arguments_classes/s2s_arguments.py |
Arguments dataclass defining all CLI configuration |
src/speech_to_speech/VAD/vad_iterator.py |
Voice-activity detection driving when to process audio |
src/speech_to_speech/STT/whisper_handler.py |
Whisper STT implementation with streaming support |
src/speech_to_speech/LLM/responses_api_language_model.py |
OpenAI Realtime API-compatible LLM client |
src/speech_to_speech/TTS/pocket_tts_handler.py |
Pocket-TTS synthesis handler |
src/speech_to_speech/connections/local_audio_streamer.py |
Microphone input capture |
src/speech_to_speech/connections/websocket_streamer.py |
WebSocket server for remote clients |
All files are available in the HuggingFace speech-to-speech repository.
Customization and Extension
Modifying Pipeline Behavior
The S2SPipeline class exposes configuration through Arguments objects. For programmatic control, import and subclass:
from src.speech_to_speech.s2s_pipeline import S2SPipeline
from src.speech_to_speech.arguments_classes.s2s_arguments import Arguments
args = Arguments(
stt="whisper_tiny",
llm="openai",
tts="pocket_tts",
# Additional parameters control VAD thresholds, buffer sizes, etc.
)
pipeline = S2SPipeline(args)
Adding Custom Handlers
The modular design in src/speech_to_speech/*/Handler.py follows a consistent pattern: implement setup(), process(audio_chunk), and warmup() methods. New backends integrate by registering in the appropriate arguments_classes module.
Troubleshooting First Runs
- No audio input: Verify microphone permissions; check
local_audio_streamer.pylogs for device enumeration - STT delays: The
vad_iterator.pyspeech threshold may need adjustment for noisy environments - LLM errors: Confirm
OPENAI_API_KEYis exported and valid for your chosen model - TTS drops: Some handlers require specific model downloads on first run; check console output for download progress
Summary
- Install with
pip install "speech-to-speech[all]"and export API keys before running - Execute the minimal demo:
python -m scripts.listen_and_play --stt whisper_tiny --llm openai --tts pocket_tts - Swap backends instantly using
--stt,--llm, and--ttsflags without code changes - Understand the flow: Audio → VAD (
vad_iterator.py) → STT → LLM (*_language_model.py) → TTS (*_handler.py) → Playback - Extend by implementing handler interfaces defined in
src/speech_to_speech/submodules
Frequently Asked Questions
What hardware requirements are needed for real-time speech-to-speech?
A modern CPU handles the whisper_tiny/pocket_tts combination comfortably. GPU acceleration significantly improves latency for larger Whisper models and neural TTS backends like kokoro or chattts. The VAD and lightweight handlers run efficiently on CPU-only systems.
Can I use self-hosted or local models instead of OpenAI APIs?
Yes. The arguments_classes system supports any OpenAI-compatible endpoint—set a custom base URL for local LLMs. For fully offline operation, combine local Whisper STT with a local LLM backend and kokoro or chattts for TTS, bypassing external API calls entirely.
How does the system handle interruptions or barge-in?
The S2SPipeline in s2s_pipeline.py manages state through the VAD iterator. When new speech is detected during TTS playback, the pipeline cancels the current synthesis cycle and begins processing the new utterance. This behavior is configurable through VAD sensitivity parameters in your Arguments.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →