How to Run the Speech-to-Speech Pipeline in Docker with GPU Support
Deploy the Hugging Face speech-to-speech pipeline with full GPU acceleration using Docker Compose and the NVIDIA Container Toolkit.
The huggingface/speech-to-speech repository provides a modular, low-latency voice conversation system that chains Voice Activity Detection (VAD), Speech-to-Text (STT), Large Language Model (LLM) inference, and Text-to-Speech (TTS) into a single streaming pipeline. Running this pipeline inside Docker with GPU support ensures that all compute-intensive stages—from audio transcription to neural speech synthesis—execute with hardware acceleration while maintaining a reproducible deployment environment.
Prerequisites: NVIDIA Container Toolkit
Before launching the containers, you must install the NVIDIA Container Toolkit on your host machine. This software enables Docker to expose GPU devices to containers via the --gpus flag (or Compose equivalents).
Follow the official installation guide at docs.nvidia.com for your distribution. Verify the installation by running nvidia-smi inside a test container:
docker run --rm --gpus all nvidia/cuda:12.0-base nvidia-smi
Understanding the Docker Architecture
The repository orchestrates two distinct services via docker-compose.yml:
llama– A GPU-enabled llama.cpp server (ghcr.io/ggml-org/llama.cpp:server-cuda) that provides local LLM inference via an OpenAI-compatible HTTP API.pipeline– The main speech-to-speech application built from the repository’sDockerfile, configured to run in socket mode and communicate with thellamaservice.
Both services declare GPU reservations using Docker’s deploy.resources.reservations.devices stanza:
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: 1
capabilities: [gpu]
This configuration grants both containers access to the host GPU without requiring manual --gpus flags at runtime.
Step-by-Step Deployment
1. Clone the Repository
Download the source code and navigate to the project root:
git clone https://github.com/huggingface/speech-to-speech.git
cd speech-to-speech
The project structure includes src/speech_to_speech/ (core pipeline logic), Dockerfile (CUDA runtime environment), and docker-compose.yml (service orchestration).
2. Build and Start the Services
Execute Docker Compose to pull the llama.cpp CUDA image and build the pipeline image:
docker compose up --build
During the build process, the Dockerfile performs the following actions:
- Uses
nvidia/cuda:12.8.1-cudnn-runtime-ubuntu24.04as the base image - Installs system dependencies (portaudio, ffmpeg, etc.)
- Creates a Python virtual environment and synchronizes dependencies via
uv sync
The pipeline service automatically starts with --mode socket and routes LLM requests to http://llama:8080/v1 as defined in the command section of docker-compose.yml.
3. Connect a Client
By default, the pipeline exposes two TCP ports for raw audio streaming:
- Port 12345 – Audio input (PCM 16 kHz, int16, mono)
- Port 12346 – Audio output
Use the provided client script to stream microphone audio and play the generated response:
python scripts/listen_and_play.py --host localhost
This script handles audio capture, transmission to the pipeline container, and playback of the synthesized speech returned on port 12346.
Configuration Options
Override the Dockerfile
To target a different architecture (e.g., ARM64), set the DOCKERFILE environment variable before building:
DOCKERFILE=Dockerfile.arm64 docker compose up --build
This variable selects an alternative build context while preserving the GPU reservation logic in docker-compose.yml.
Switch to Realtime Mode
The default configuration runs the pipeline in socket mode. To use the WebSocket-based OpenAI Realtime API instead, modify the pipeline service command in docker-compose.yml:
command: ["--mode", "realtime", "--llm-backend", "openai-realtime"]
The GPU configuration remains identical regardless of the operating mode.
Summary
- huggingface/speech-to-speech implements a four-stage pipeline (VAD → STT → LLM → TTS) where each component executes in its own thread and communicates via queues.
- GPU acceleration requires the NVIDIA Container Toolkit and explicit device reservations in
docker-compose.ymlusing thedeploy.resources.reservations.devicesstanza. - The
llamaservice runs llama.cpp with CUDA support, while thepipelineservice processes audio I/O via TCP sockets on ports 12345 and 12346. - Build and launch the entire stack with
docker compose up --build, then connect usingscripts/listen_and_play.py.
Frequently Asked Questions
What NVIDIA driver version is required?
You need a driver supporting CUDA 12.8 or later, as the Dockerfile bases the pipeline container on nvidia/cuda:12.8.1-cudnn-runtime-ubuntu24.04. Verify compatibility by running nvidia-smi on the host and confirming the CUDA Version is 12.0 or higher.
Can I run the pipeline without a GPU?
Yes, but you must modify the docker-compose.yml to remove the GPU reservation stanzas and change the llama service image to a CPU-only variant (e.g., ghcr.io/ggml-org/llama.cpp:server). Performance will degrade significantly for the LLM and TTS stages.
How do I change the default STT or TTS models?
Mount a custom configuration file or override environment variables in docker-compose.yml. The pipeline loads model specifications from arguments passed to the speech-to-speech CLI entry point. Refer to src/speech_to_speech/ module files to identify the exact parameter names for Parakeet TDT (STT) and Qwen3-TTS (TTS) backends.
Why does the container need both the llama service and the pipeline service?
The llama service provides a dedicated, optimized inference endpoint for the LLM stage using llama.cpp, while the pipeline service handles the VAD, STT, and TTS components. This separation allows independent scaling and permits the pipeline to swap between local LLMs (llama.cpp) and remote APIs (OpenAI) without rebuilding the container image.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →