How to Run a Basic Speech-to-Speech Example with HuggingFace: Complete Setup Guide

Run a real-time speech-to-speech pipeline using python -m scripts.listen_and_play --stt whisper_tiny --llm openai --tts pocket_tts after installing the package and setting your API keys.

The HuggingFace speech-to-speech repository provides a modular, real-time pipeline that chains together speech recognition (STT), language generation (LLM), and speech synthesis (TTS) into a single streaming system. This guide walks you through running the minimal "listen-and-play" demo that captures your microphone audio, processes it through all three stages, and plays back the synthesized response.

Understanding the Pipeline Architecture

The speech-to-speech system operates as a coordinated chain of specialized handlers orchestrated by S2SPipeline in src/speech_to_speech/s2s_pipeline.py. Understanding this flow helps you troubleshoot and customize your setup.

Data Flow Through the System

  1. Audio Capture — local_audio_streamer.py reads microphone input continuously
  2. Voice Detection — vad_iterator.py identifies speech segments to process
  3. Speech-to-Text — One of STT/*_handler.py modules (Whisper, Paraformer, MMS) transcribes audio
  4. Language Generation — An LLM backend (responses_api_language_model.py, chat_completions_language_model.py) generates responses
  5. Text-to-Speech — A TTS handler (pocket_tts_handler.py, qwen3_tts_handler.py, etc.) synthesizes audio
  6. Playback — listen_and_play.py sends audio directly to your speakers or streams via websocket_streamer.py

All components share a unified configuration through typed Arguments classes defined in src/speech_to_speech/arguments_classes/*_arguments.py, making backend swaps as simple as changing a CLI flag.

Installation and Prerequisites

Step 1: Install the Package

The repository supports optional dependencies for different model backends. For the simplest setup, install with all extras:

pip install "speech-to-speech[all]"

For a lighter installation, you can specify only the backends you need.

Step 2: Configure API Credentials

Most LLM backends require authentication. Export your keys before running:

export OPENAI_API_KEY=your_openai_key_here

Other backends may need additional environment variables—check the handler's argument class for specifics.

Running Your First Speech-to-Speech Demo

The Minimal Command

Execute the end-to-end pipeline with three CLI flags selecting your backends:

python -m scripts.listen_and_play \
    --stt whisper_tiny \
    --llm openai \
    --tts pocket_tts

What happens when you run this:

Speak into your microphone—you'll hear your synthesized response within seconds.

Alternative Backend Combinations

Swap any component by changing its flag. The pipeline automatically loads the corresponding handler from src/speech_to_speech/:


# Use Paraformer for STT, Responses API for LLM, Qwen-3 for TTS

python -m scripts.listen_and_play \
    --stt paraformer \
    --llm responses_api \
    --tts qwen3

Available options are defined in the argument classes and include:

  • STT: whisper_tiny, whisper_base, whisper_large, paraformer, mms
  • LLM: openai, responses_api, chat_completions (plus any OpenAI-compatible endpoint)
  • TTS: pocket_tts, qwen3, kokoro, chattts

Key Source Files for Reference

File Path Purpose
scripts/listen_and_play.py Entry-point demo script that wires all components together
src/speech_to_speech/s2s_pipeline.py Core orchestration; creates handlers and manages the processing loop
src/speech_to_speech/arguments_classes/s2s_arguments.py Arguments dataclass defining all CLI configuration
src/speech_to_speech/VAD/vad_iterator.py Voice-activity detection driving when to process audio
src/speech_to_speech/STT/whisper_handler.py Whisper STT implementation with streaming support
src/speech_to_speech/LLM/responses_api_language_model.py OpenAI Realtime API-compatible LLM client
src/speech_to_speech/TTS/pocket_tts_handler.py Pocket-TTS synthesis handler
src/speech_to_speech/connections/local_audio_streamer.py Microphone input capture
src/speech_to_speech/connections/websocket_streamer.py WebSocket server for remote clients

All files are available in the HuggingFace speech-to-speech repository.

Customization and Extension

Modifying Pipeline Behavior

The S2SPipeline class exposes configuration through Arguments objects. For programmatic control, import and subclass:

from src.speech_to_speech.s2s_pipeline import S2SPipeline
from src.speech_to_speech.arguments_classes.s2s_arguments import Arguments

args = Arguments(
    stt="whisper_tiny",
    llm="openai",
    tts="pocket_tts",
    # Additional parameters control VAD thresholds, buffer sizes, etc.

)
pipeline = S2SPipeline(args)

Adding Custom Handlers

The modular design in src/speech_to_speech/*/Handler.py follows a consistent pattern: implement setup(), process(audio_chunk), and warmup() methods. New backends integrate by registering in the appropriate arguments_classes module.

Troubleshooting First Runs

  • No audio input: Verify microphone permissions; check local_audio_streamer.py logs for device enumeration
  • STT delays: The vad_iterator.py speech threshold may need adjustment for noisy environments
  • LLM errors: Confirm OPENAI_API_KEY is exported and valid for your chosen model
  • TTS drops: Some handlers require specific model downloads on first run; check console output for download progress

Summary

  • Install with pip install "speech-to-speech[all]" and export API keys before running
  • Execute the minimal demo: python -m scripts.listen_and_play --stt whisper_tiny --llm openai --tts pocket_tts
  • Swap backends instantly using --stt, --llm, and --tts flags without code changes
  • Understand the flow: Audio → VAD (vad_iterator.py) → STT → LLM (*_language_model.py) → TTS (*_handler.py) → Playback
  • Extend by implementing handler interfaces defined in src/speech_to_speech/ submodules

Frequently Asked Questions

What hardware requirements are needed for real-time speech-to-speech?

A modern CPU handles the whisper_tiny/pocket_tts combination comfortably. GPU acceleration significantly improves latency for larger Whisper models and neural TTS backends like kokoro or chattts. The VAD and lightweight handlers run efficiently on CPU-only systems.

Can I use self-hosted or local models instead of OpenAI APIs?

Yes. The arguments_classes system supports any OpenAI-compatible endpoint—set a custom base URL for local LLMs. For fully offline operation, combine local Whisper STT with a local LLM backend and kokoro or chattts for TTS, bypassing external API calls entirely.

How does the system handle interruptions or barge-in?

The S2SPipeline in s2s_pipeline.py manages state through the VAD iterator. When new speech is detected during TTS playback, the pipeline cancels the current synthesis cycle and begins processing the new utterance. This behavior is configurable through VAD sensitivity parameters in your Arguments.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →