# How to Run a Basic Speech-to-Speech Example with HuggingFace: Complete Setup Guide

> Learn how to run a basic speech-to-speech example with HuggingFace. Follow our complete setup guide to get your real-time pipeline running quickly.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: getting-started
- Published: 2026-08-02

---

**Run a real-time speech-to-speech pipeline using `python -m scripts.listen_and_play --stt whisper_tiny --llm openai --tts pocket_tts` after installing the package and setting your API keys.**

The **HuggingFace speech-to-speech** repository provides a modular, real-time pipeline that chains together speech recognition (STT), language generation (LLM), and speech synthesis (TTS) into a single streaming system. This guide walks you through running the minimal "listen-and-play" demo that captures your microphone audio, processes it through all three stages, and plays back the synthesized response.

## Understanding the Pipeline Architecture

The speech-to-speech system operates as a coordinated chain of specialized handlers orchestrated by `S2SPipeline` in [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py). Understanding this flow helps you troubleshoot and customize your setup.

### Data Flow Through the System

1. **Audio Capture** — [`local_audio_streamer.py`](https://github.com/huggingface/speech-to-speech/blob/main/local_audio_streamer.py) reads microphone input continuously
2. **Voice Detection** — [`vad_iterator.py`](https://github.com/huggingface/speech-to-speech/blob/main/vad_iterator.py) identifies speech segments to process
3. **Speech-to-Text** — One of `STT/*_handler.py` modules (Whisper, Paraformer, MMS) transcribes audio
4. **Language Generation** — An LLM backend ([`responses_api_language_model.py`](https://github.com/huggingface/speech-to-speech/blob/main/responses_api_language_model.py), [`chat_completions_language_model.py`](https://github.com/huggingface/speech-to-speech/blob/main/chat_completions_language_model.py)) generates responses
5. **Text-to-Speech** — A TTS handler ([`pocket_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/pocket_tts_handler.py), [`qwen3_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/qwen3_tts_handler.py), etc.) synthesizes audio
6. **Playback** — [`listen_and_play.py`](https://github.com/huggingface/speech-to-speech/blob/main/listen_and_play.py) sends audio directly to your speakers or streams via [`websocket_streamer.py`](https://github.com/huggingface/speech-to-speech/blob/main/websocket_streamer.py)

All components share a unified configuration through typed `Arguments` classes defined in `src/speech_to_speech/arguments_classes/*_arguments.py`, making backend swaps as simple as changing a CLI flag.

## Installation and Prerequisites

### Step 1: Install the Package

The repository supports optional dependencies for different model backends. For the simplest setup, install with all extras:

```bash
pip install "speech-to-speech[all]"

```

For a lighter installation, you can specify only the backends you need.

### Step 2: Configure API Credentials

Most LLM backends require authentication. Export your keys before running:

```bash
export OPENAI_API_KEY=your_openai_key_here

```

Other backends may need additional environment variables—check the handler's argument class for specifics.

## Running Your First Speech-to-Speech Demo

### The Minimal Command

Execute the end-to-end pipeline with three CLI flags selecting your backends:

```bash
python -m scripts.listen_and_play \
    --stt whisper_tiny \
    --llm openai \
    --tts pocket_tts

```

**What happens when you run this:**

- [`scripts/listen_and_play.py`](https://github.com/huggingface/speech-to-speech/blob/main/scripts/listen_and_play.py) instantiates `S2SPipeline` with `Arguments` parsed from your flags
- The pipeline creates handlers: [`whisper_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/whisper_handler.py) for STT, an OpenAI-compatible LLM client, and [`pocket_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/pocket_tts_handler.py) for synthesis
- [`local_audio_streamer.py`](https://github.com/huggingface/speech-to-speech/blob/main/local_audio_streamer.py) begins capturing microphone audio through the VAD
- Each detected speech segment flows through STT → LLM → TTS → local audio output

Speak into your microphone—you'll hear your synthesized response within seconds.

### Alternative Backend Combinations

Swap any component by changing its flag. The pipeline automatically loads the corresponding handler from `src/speech_to_speech/`:

```bash

# Use Paraformer for STT, Responses API for LLM, Qwen-3 for TTS

python -m scripts.listen_and_play \
    --stt paraformer \
    --llm responses_api \
    --tts qwen3

```

Available options are defined in the argument classes and include:

- **STT:** `whisper_tiny`, `whisper_base`, `whisper_large`, `paraformer`, `mms`
- **LLM:** `openai`, `responses_api`, `chat_completions` (plus any OpenAI-compatible endpoint)
- **TTS:** `pocket_tts`, `qwen3`, `kokoro`, `chattts`

## Key Source Files for Reference

| File Path | Purpose |
|-----------|---------|
| [`scripts/listen_and_play.py`](https://github.com/huggingface/speech-to-speech/blob/main/scripts/listen_and_play.py) | Entry-point demo script that wires all components together |
| [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py) | Core orchestration; creates handlers and manages the processing loop |
| [`src/speech_to_speech/arguments_classes/s2s_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/s2s_arguments.py) | `Arguments` dataclass defining all CLI configuration |
| [`src/speech_to_speech/VAD/vad_iterator.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/VAD/vad_iterator.py) | Voice-activity detection driving when to process audio |
| [`src/speech_to_speech/STT/whisper_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/STT/whisper_handler.py) | Whisper STT implementation with streaming support |
| [`src/speech_to_speech/LLM/responses_api_language_model.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/responses_api_language_model.py) | OpenAI Realtime API-compatible LLM client |
| [`src/speech_to_speech/TTS/pocket_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/pocket_tts_handler.py) | Pocket-TTS synthesis handler |
| [`src/speech_to_speech/connections/local_audio_streamer.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/connections/local_audio_streamer.py) | Microphone input capture |
| [`src/speech_to_speech/connections/websocket_streamer.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/connections/websocket_streamer.py) | WebSocket server for remote clients |

All files are available in the [HuggingFace speech-to-speech repository](https://github.com/huggingface/speech-to-speech/tree/main).

## Customization and Extension

### Modifying Pipeline Behavior

The `S2SPipeline` class exposes configuration through `Arguments` objects. For programmatic control, import and subclass:

```python
from src.speech_to_speech.s2s_pipeline import S2SPipeline
from src.speech_to_speech.arguments_classes.s2s_arguments import Arguments

args = Arguments(
    stt="whisper_tiny",
    llm="openai",
    tts="pocket_tts",
    # Additional parameters control VAD thresholds, buffer sizes, etc.

)
pipeline = S2SPipeline(args)

```

### Adding Custom Handlers

The modular design in `src/speech_to_speech/*/Handler.py` follows a consistent pattern: implement `setup()`, `process(audio_chunk)`, and `warmup()` methods. New backends integrate by registering in the appropriate `arguments_classes` module.

## Troubleshooting First Runs

- **No audio input:** Verify microphone permissions; check [`local_audio_streamer.py`](https://github.com/huggingface/speech-to-speech/blob/main/local_audio_streamer.py) logs for device enumeration
- **STT delays:** The [`vad_iterator.py`](https://github.com/huggingface/speech-to-speech/blob/main/vad_iterator.py) speech threshold may need adjustment for noisy environments
- **LLM errors:** Confirm `OPENAI_API_KEY` is exported and valid for your chosen model
- **TTS drops:** Some handlers require specific model downloads on first run; check console output for download progress

## Summary

- **Install** with `pip install "speech-to-speech[all]"` and **export API keys** before running
- **Execute** the minimal demo: `python -m scripts.listen_and_play --stt whisper_tiny --llm openai --tts pocket_tts`
- **Swap backends** instantly using `--stt`, `--llm`, and `--tts` flags without code changes
- **Understand the flow:** Audio → VAD ([`vad_iterator.py`](https://github.com/huggingface/speech-to-speech/blob/main/vad_iterator.py)) → STT → LLM (`*_language_model.py`) → TTS (`*_handler.py`) → Playback
- **Extend** by implementing handler interfaces defined in `src/speech_to_speech/` submodules

## Frequently Asked Questions

### What hardware requirements are needed for real-time speech-to-speech?

A modern CPU handles the `whisper_tiny`/`pocket_tts` combination comfortably. GPU acceleration significantly improves latency for larger Whisper models and neural TTS backends like `kokoro` or `chattts`. The VAD and lightweight handlers run efficiently on CPU-only systems.

### Can I use self-hosted or local models instead of OpenAI APIs?

Yes. The `arguments_classes` system supports any OpenAI-compatible endpoint—set a custom base URL for local LLMs. For fully offline operation, combine local Whisper STT with a local LLM backend and `kokoro` or `chattts` for TTS, bypassing external API calls entirely.

### How does the system handle interruptions or barge-in?

The `S2SPipeline` in [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py) manages state through the VAD iterator. When new speech is detected during TTS playback, the pipeline cancels the current synthesis cycle and begins processing the new utterance. This behavior is configurable through VAD sensitivity parameters in your `Arguments`.