Setting Up Docker Deployment with GPU Passthrough for Production
To deploy the Hugging Face Speech-to-Speech pipeline with GPU passthrough in production, configure the docker-compose.yml to reserve NVIDIA devices via the deploy.resources.reservations.devices stanza, enabling both the LLM backend and pipeline containers to access CUDA drivers simultaneously.
The huggingface/speech-to-speech repository ships a production-ready containerization strategy designed for NVIDIA GPU acceleration. Setting up Docker deployment with GPU passthrough for production requires configuring the provided Dockerfile and docker-compose.yml to ensure the NVIDIA Container Toolkit can inject driver libraries into the runtime environment.
Prerequisites for NVIDIA GPU Passthrough
Before deploying containers, install the NVIDIA Container Toolkit on the host machine. This toolkit is essential because it intercepts Docker container creation requests and injects the necessary driver libraries, creating the /dev/nvidia* devices inside the container. Without this toolkit, the deploy.resources.reservations.devices configuration in Docker Compose will fail to allocate GPU resources.
Core Architecture and Configuration Files
The deployment architecture centers on two files in the repository root that work together to provide CUDA support:
Dockerfile CUDA Base Image
The Dockerfile builds upon nvidia/cuda:12.8.1-cudnn-runtime-ubuntu24.04, providing CUDA 12.8, cuDNN, and the runtime libraries required by deep learning frameworks like PyTorch and Transformers. The build process installs system dependencies, creates a virtual environment, and uses uv (a fast Python package resolver) to install the speech-to-speech package without development dependencies. This ensures all Python extensions are compiled against the specific CUDA version present in the image.
Docker Compose Orchestration
The docker-compose.yml defines two interconnected services that both require GPU access:
| Component | Role | GPU Interaction |
|---|---|---|
| llama | Runs the pre-built ggml-org/llama.cpp:server-cuda image |
Executes a CUDA-accelerated binary for GGUF model inference |
| pipeline | Builds from the repository's Dockerfile |
Runs speech-to-speech serve with access to the reserved GPU |
Both services use the deploy.resources.reservations.devices field to request GPU resources, ensuring the host's NVIDIA driver is accessible inside each container.
Configuring GPU Resources in Docker Compose
The GPU passthrough mechanism relies on the deploy block within each service definition. The configuration specifies driver: nvidia and capabilities: [gpu], which instructs Docker to expose the host's GPU to the container. As implemented in huggingface/speech-to-speech, this allows the llama service and the pipeline service to share GPU 0 without conflicting over device access.
Step-by-Step Deployment Guide
Building and Starting the Production Stack
To build the custom image and start both services with GPU support, run:
DOCKERFILE=Dockerfile docker compose up -d
This command builds the pipeline image using the specified Dockerfile and starts the orchestrated stack. The -d flag runs containers in detached mode suitable for production environments.
Running a Standalone Container for Testing
For local testing without Docker Compose, build and run the container directly:
docker build -t s2s-demo -f Dockerfile .
docker run --gpus all -p 8765:8765 s2s-demo \
speech-to-speech serve \
--stt whisper \
--tts pocket \
--host 0.0.0.0 \
--port 8765
The --gpus all flag tells Docker to pass through all available NVIDIA GPUs to the container.
Verifying GPU Passthrough
Confirm that the containers can access the host GPU by executing nvidia-smi inside the running pipeline container:
docker exec -it $(docker ps -qf "name=pipeline") nvidia-smi
The output should display the host GPU with driver versions matching the host system, confirming that the NVIDIA Container Toolkit successfully injected the drivers.
Low-Latency GPU Configuration
For production environments requiring minimal latency, the repository supports co-locating the TTS, STT, and LLM processing on the same GPU. According to src/speech_to_speech/TTS/README.md, the --device argument (defined in src/speech_to_speech/arguments_classes/*_arguments.py) defaults to cuda. When the container has GPU access via the reservation configuration, this default setting enables immediate GPU acceleration without additional flags.
Summary
- The
Dockerfileusesnvidia/cuda:12.8.1-cudnn-runtime-ubuntu24.04to ensure CUDA 12.8 compatibility and includes cuDNN runtime libraries docker-compose.ymlorchestrates thellamaandpipelineservices with explicit GPU reservations usingdeploy.resources.reservations.devices- The NVIDIA Container Toolkit must be installed on the host to enable driver injection and
/dev/nvidia*device creation inside containers - Verify successful GPU passthrough by running
nvidia-smiinside the running container - The
speech-to-speech servecommand uses--device cudaby default when GPU resources are available
Frequently Asked Questions
Do I need to install CUDA drivers inside the container?
No. The Dockerfile inherits from the official NVIDIA CUDA runtime base image, which includes the necessary CUDA libraries and cuDNN. The NVIDIA Container Toolkit handles binding the host's NVIDIA drivers into the container at runtime, so you do not need to install drivers manually.
How do I specify which GPU to use if the host has multiple cards?
The default docker-compose.yml configuration reserves GPU 0. You can modify the device_ids field under deploy.resources.reservations.devices to specify alternative GPU indices (for example, "1" for the second GPU) or use "all" to pass through every available GPU to the container.
Can I deploy the pipeline without the included LLM service?
Yes. While the default docker-compose.yml includes the llama service for convenience, you can run the pipeline container standalone using docker run --gpus all as shown in the deployment steps. Alternatively, modify the compose file to remove the llama service dependency if you are using an external LLM backend API.
Why does the build process use uv instead of pip?
The Dockerfile utilizes uv for faster Python package resolution and installation. This approach installs the speech-to-speech package and its dependencies more quickly than traditional pip, excludes development dependencies to reduce image size, and ensures the installed wheels are compiled specifically for the CUDA version present in the base image.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →