How to Set Up Docker Deployment for Production with Fish Speech
Fish Speech provides a multi-stage Docker build system that produces production-ready images for both interactive web interfaces and REST API servers, supporting both CUDA and CPU backends with integrated health checks.
Setting up Docker deployment for production with the fishaudio/fish-speech repository requires understanding its backend-agnostic build architecture. The project provides distinct build targets for different deployment scenarios, allowing you to optimize container images for either the Gradio-based web UI or the FastAPI server implementation.
Understanding the Multi-Stage Docker Architecture
The production Dockerfile at docker/Dockerfile implements a classic multi-stage build pattern that minimizes image size while maximizing flexibility. This architecture separates backend selection, shared dependencies, and target-specific configurations into distinct stages.
Backend Selection and Base Images
The build process begins with a BACKEND build argument that defaults to cuda but can be switched to cpu for environments without NVIDIA GPUs (lines 63-67).
base-cuda: Uses the NVIDIA CUDA runtime image for GPU-accelerated inferencebase-cpu: Uses a slim Python base image for lightweight CPU-only deployments
Shared Application Layer
The app stage installs system audio dependencies, pulls the uv Python installer for fast package management, creates a non-root user for security, and copies the source code. This layer is reused by both target stages to ensure consistency.
Target-Specific Stages
The Dockerfile defines two production targets:
webui: Configures Gradio-specific environment variables (GRADIO_SERVER_NAME,GRADIO_SERVER_PORT) and generates a startup script that executestools/run_webui.pyserver: Sets API server environment variables and creates a startup script fortools/api_server.py
Both targets include built-in health checks at lines 222-224 (webui) and lines 370-372 (server), exposing ports 7860 and 8080 respectively.
Building Production Images
To create a production image for the web interface with CUDA support:
docker build \
--platform linux/amd64 \
-f docker/Dockerfile \
--build-arg BACKEND=cuda \
--target webui \
-t fish-speech-webui:cuda .
For the API server with CPU backend:
docker build \
--platform linux/amd64 \
-f docker/Dockerfile \
--build-arg BACKEND=cpu \
--target server \
-t fish-speech-server:cpu .
Replace --target webui with --target server to switch between interfaces, and adjust BACKEND to match your hardware capabilities.
Preparing Model Checkpoints
The Docker images do not contain model weights. You must mount a local checkpoints directory into the container at /app/checkpoints:
# Example directory structure
checkpoints/
└── s2-pro/
├── codec.pth
└── ... (other model files)
Use Docker volumes or bind mounts to persist these files outside the container for production deployments.
Running Containers in Production
Web UI Deployment
For GPU-enabled production deployment of the web interface:
docker run -d \
--gpus all \
-v $(pwd)/checkpoints:/app/checkpoints \
-e COMPILE=1 \
-p 7860:7860 \
--name fish-speech-webui \
fish-speech-webui:cuda
The COMPILE=1 environment variable forces on-the-fly model compilation for optimized inference performance.
API Server Deployment
For the REST API server:
docker run -d \
--gpus all \
-v $(pwd)/checkpoints:/app/checkpoints \
-e COMPILE=0 \
-p 8080:8080 \
--name fish-speech-server \
fish-speech-server:cuda
Setting COMPILE=0 skips compilation for faster container startup, trading initial latency for runtime performance.
CPU-Only Deployment
For environments without NVIDIA GPUs, omit the --gpus all flag and use the CPU-tagged image:
docker run -d \
-v $(pwd)/checkpoints:/app/checkpoints \
-p 8080:8080 \
--name fish-speech-server-cpu \
fish-speech-server:cpu
Health Checks and Monitoring
Both production images include built-in health check endpoints for load balancer integration and monitoring systems:
# Verify Web UI health
curl -f http://localhost:7860/health && echo "UI healthy"
# Verify API Server health
curl -f http://localhost:8080/v1/health && echo "API healthy"
The health check configurations are defined at lines 222-224 for the webui target and lines 370-372 for the server target in the Dockerfile.
Key Files and Entry Points
| File | Role | Location |
|---|---|---|
docker/Dockerfile |
Multi-stage build definition with backend selection and target-specific stages | docker/Dockerfile |
tools/run_webui.py |
Gradio interface entry point invoked by webui target | tools/run_webui.py |
tools/api_server.py |
FastAPI server entry point invoked by server target | tools/api_server.py |
docker/common.sh |
Runtime helper functions for logging, device validation, and health checks | Generated at build time (see Dockerfile creation section) |
Summary
- Fish Speech provides a multi-stage Dockerfile at
docker/Dockerfilesupporting bothwebuiandserverproduction targets. - Backend flexibility allows building for
cuda(GPU) orcpu(CPU-only) via theBACKENDbuild argument. - Model checkpoints must be mounted externally to
/app/checkpointsas the images do not contain model weights. - Health checks are built into both targets at
/health(webui) and/v1/health(server) for production monitoring. - Compilation control via the
COMPILEenvironment variable allows trading startup time for runtime performance optimization.
Frequently Asked Questions
How do I switch between CUDA and CPU backends?
Set the BACKEND build argument to cpu during the Docker build process. The Dockerfile at lines 63-67 handles backend selection, switching the base image from NVIDIA CUDA runtime to a slim Python image. For runtime execution, CPU-only containers do not require the --gpus all Docker flag.
Where should I mount model checkpoints?
Mount your local checkpoints directory to /app/checkpoints inside the container using the -v flag. The Docker images do not include model weights, so this mount is mandatory for operation. Ensure the directory contains the appropriate subdirectories (such as s2-pro/ with codec.pth and related files) before starting the container.
What is the difference between the webui and server targets?
The webui target builds a Gradio-based interactive interface accessible via browser on port 7860, while the server target builds a FastAPI REST API service on port 8080. The webui target executes tools/run_webui.py and sets Gradio-specific environment variables, whereas the server target executes tools/api_server.py with API-specific configurations. Both include health checks but expose different endpoints (/health vs /v1/health).
How do I enable model compilation for faster inference?
Set the environment variable COMPILE=1 when running the container. This forces on-the-fly model compilation using the entrypoint logic defined in the Dockerfile at lines 35-46. Note that compilation requires write permissions to /app/references inside the container, so ensure this directory is mounted with appropriate permissions. For faster container startup at the expense of initial request latency, use COMPILE=0 instead.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →