# How to Run the Speech-to-Speech Pipeline in Docker with GPU Support

> Easily run the speech-to-speech pipeline in Docker with GPU acceleration. Deploy the Hugging Face model quickly using Docker Compose and NVIDIA Container Toolkit for optimal performance.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: how-to-guide
- Published: 2026-07-30

---

**Deploy the Hugging Face speech-to-speech pipeline with full GPU acceleration using Docker Compose and the NVIDIA Container Toolkit.**

The **huggingface/speech-to-speech** repository provides a modular, low-latency voice conversation system that chains Voice Activity Detection (VAD), Speech-to-Text (STT), Large Language Model (LLM) inference, and Text-to-Speech (TTS) into a single streaming pipeline. Running this pipeline inside Docker with GPU support ensures that all compute-intensive stages—from audio transcription to neural speech synthesis—execute with hardware acceleration while maintaining a reproducible deployment environment.

## Prerequisites: NVIDIA Container Toolkit

Before launching the containers, you must install the **NVIDIA Container Toolkit** on your host machine. This software enables Docker to expose GPU devices to containers via the `--gpus` flag (or Compose equivalents).

Follow the official installation guide at [docs.nvidia.com](https://docs.nvidia.com/datacenter/cloud-native/container-toolkit/latest/install-guide.html) for your distribution. Verify the installation by running `nvidia-smi` inside a test container:

```bash
docker run --rm --gpus all nvidia/cuda:12.0-base nvidia-smi

```

## Understanding the Docker Architecture

The repository orchestrates two distinct services via [`docker-compose.yml`](https://github.com/huggingface/speech-to-speech/blob/main/docker-compose.yml):

1. **`llama`** – A GPU-enabled **llama.cpp** server (`ghcr.io/ggml-org/llama.cpp:server-cuda`) that provides local LLM inference via an OpenAI-compatible HTTP API.
2. **`pipeline`** – The main speech-to-speech application built from the repository’s `Dockerfile`, configured to run in socket mode and communicate with the `llama` service.

Both services declare GPU reservations using Docker’s `deploy.resources.reservations.devices` stanza:

```yaml
deploy:
  resources:
    reservations:
      devices:
        - driver: nvidia
          count: 1
          capabilities: [gpu]

```

This configuration grants both containers access to the host GPU without requiring manual `--gpus` flags at runtime.

## Step-by-Step Deployment

### 1. Clone the Repository

Download the source code and navigate to the project root:

```bash
git clone https://github.com/huggingface/speech-to-speech.git
cd speech-to-speech

```

The project structure includes `src/speech_to_speech/` (core pipeline logic), `Dockerfile` (CUDA runtime environment), and [`docker-compose.yml`](https://github.com/huggingface/speech-to-speech/blob/main/docker-compose.yml) (service orchestration).

### 2. Build and Start the Services

Execute Docker Compose to pull the llama.cpp CUDA image and build the pipeline image:

```bash
docker compose up --build

```

During the build process, the `Dockerfile` performs the following actions:
- Uses `nvidia/cuda:12.8.1-cudnn-runtime-ubuntu24.04` as the base image
- Installs system dependencies (portaudio, ffmpeg, etc.)
- Creates a Python virtual environment and synchronizes dependencies via `uv sync`

The `pipeline` service automatically starts with `--mode socket` and routes LLM requests to `http://llama:8080/v1` as defined in the `command` section of [`docker-compose.yml`](https://github.com/huggingface/speech-to-speech/blob/main/docker-compose.yml).

### 3. Connect a Client

By default, the pipeline exposes two TCP ports for raw audio streaming:
- **Port 12345** – Audio input (PCM 16 kHz, int16, mono)
- **Port 12346** – Audio output

Use the provided client script to stream microphone audio and play the generated response:

```bash
python scripts/listen_and_play.py --host localhost

```

This script handles audio capture, transmission to the pipeline container, and playback of the synthesized speech returned on port 12346.

## Configuration Options

### Override the Dockerfile

To target a different architecture (e.g., ARM64), set the `DOCKERFILE` environment variable before building:

```bash
DOCKERFILE=Dockerfile.arm64 docker compose up --build

```

This variable selects an alternative build context while preserving the GPU reservation logic in [`docker-compose.yml`](https://github.com/huggingface/speech-to-speech/blob/main/docker-compose.yml).

### Switch to Realtime Mode

The default configuration runs the pipeline in **socket** mode. To use the WebSocket-based OpenAI Realtime API instead, modify the `pipeline` service command in [`docker-compose.yml`](https://github.com/huggingface/speech-to-speech/blob/main/docker-compose.yml):

```yaml
command: ["--mode", "realtime", "--llm-backend", "openai-realtime"]

```

The GPU configuration remains identical regardless of the operating mode.

## Summary

- **huggingface/speech-to-speech** implements a four-stage pipeline (VAD → STT → LLM → TTS) where each component executes in its own thread and communicates via queues.
- GPU acceleration requires the **NVIDIA Container Toolkit** and explicit device reservations in [`docker-compose.yml`](https://github.com/huggingface/speech-to-speech/blob/main/docker-compose.yml) using the `deploy.resources.reservations.devices` stanza.
- The `llama` service runs **llama.cpp** with CUDA support, while the `pipeline` service processes audio I/O via TCP sockets on ports 12345 and 12346.
- Build and launch the entire stack with `docker compose up --build`, then connect using [`scripts/listen_and_play.py`](https://github.com/huggingface/speech-to-speech/blob/main/scripts/listen_and_play.py).

## Frequently Asked Questions

### What NVIDIA driver version is required?

You need a driver supporting CUDA 12.8 or later, as the `Dockerfile` bases the pipeline container on `nvidia/cuda:12.8.1-cudnn-runtime-ubuntu24.04`. Verify compatibility by running `nvidia-smi` on the host and confirming the CUDA Version is 12.0 or higher.

### Can I run the pipeline without a GPU?

Yes, but you must modify the [`docker-compose.yml`](https://github.com/huggingface/speech-to-speech/blob/main/docker-compose.yml) to remove the GPU reservation stanzas and change the `llama` service image to a CPU-only variant (e.g., `ghcr.io/ggml-org/llama.cpp:server`). Performance will degrade significantly for the LLM and TTS stages.

### How do I change the default STT or TTS models?

Mount a custom configuration file or override environment variables in [`docker-compose.yml`](https://github.com/huggingface/speech-to-speech/blob/main/docker-compose.yml). The pipeline loads model specifications from arguments passed to the `speech-to-speech` CLI entry point. Refer to `src/speech_to_speech/` module files to identify the exact parameter names for Parakeet TDT (STT) and Qwen3-TTS (TTS) backends.

### Why does the container need both the llama service and the pipeline service?

The **llama** service provides a dedicated, optimized inference endpoint for the LLM stage using **llama.cpp**, while the **pipeline** service handles the VAD, STT, and TTS components. This separation allows independent scaling and permits the pipeline to swap between local LLMs (llama.cpp) and remote APIs (OpenAI) without rebuilding the container image.