How to Deploy the VoiceStudio Backend: Docker Installation Guide

Deploy the VoiceStudio backend using the official Docker image from GitHub Container Registry by running a single docker run command with the OMNIVOICE_API_KEY environment variable and volume mounts for omnivoice-data and HuggingFace cache.

VoiceStudio's backend is packaged as a Docker container that bundles the FastAPI server (uvicorn backend.main:app), Python/PyTorch runtime, compiled React frontend, and system dependencies like FFmpeg and libsndfile. According to the debpalash/VoiceStudio repository, the container exposes port 3900 and requires specific environment variables to enable external access and secure administrative endpoints defined in backend/core/auth.py.

One-Command Docker Deployment

The fastest way to deploy VoiceStudio uses docker run with the appropriate image tag for your hardware. The container requires two persistent volumes (omnivoice-data for application data and ~/.cache/huggingface for model weights) and the OMNIVOICE_API_KEY environment variable.

First, generate a secure administrator API key:

export OMNIVOICE_API_KEY="$(python3 -c 'import secrets; print(secrets.token_urlsafe(32))')"

CPU Deployment

Pull the latest image and run on CPU:

docker pull ghcr.io/debpalash/omnivoice-studio:latest

docker run -d --name omnivoice \
  -p 127.0.0.1:3900:3900 \
  -e OMNIVOICE_API_KEY="$OMNIVOICE_API_KEY" \
  -v omnivoice-data:/app/omnivoice_data \
  -v ~/.cache/huggingface:/root/.cache/huggingface \
  ghcr.io/debpalash/omnivoice-studio:latest

NVIDIA GPU Deployment

Use the same image (CUDA-enabled) with the --gpus all flag:

docker run -d --name omnivoice --gpus all \
  -p 127.0.0.1:3900:3900 \
  -e OMNIVOICE_API_KEY="$OMNIVOICE_API_KEY" \
  -v omnivoice-data:/app/omnivoice_data \
  -v ~/.cache/huggingface:/root/.cache/huggingface \
  ghcr.io/debpalash/omnivoice-studio:latest

AMD ROCm Deployment

Pull the ROCm-specific tag and pass the GPU device nodes:

docker pull ghcr.io/debpalash/omnivoice-studio:rocm

docker run -d --name omnivoice \
  --device /dev/kfd --device /dev/dri \
  -p 127.0.0.1:3900:3900 \
  -e OMNIVOICE_API_KEY="$OMNIVOICE_API_KEY" \
  -v omnivoice-data:/app/lomnivoice_data \
  -v ~/.cache/huggingface:/root/.cache/huggingface \
  ghcr.io/debpalash/omnivoice-studio:rocm

The UI becomes available at http://localhost:3900 once the container health check passes. Initial startup downloads 2–4 GB of model weights; monitor progress with docker logs -f omnivoice.

The repository's deploy/docker-compose.yml defines profiles for CPU, GPU, and ROCm that automatically configure volumes and environment variables. This approach is defined in the repository's deployment configuration at /deploy/docker-compose.yml.

export OMNIVOICE_API_KEY="$(python3 -c 'import secrets; print(secrets.token_urlsafe(32))')"

# CPU profile

docker compose -f deploy/docker-compose.yml --profile cpu up -d

# NVIDIA GPU profile

docker compose -f deploy/docker-compose.yml --profile gpu up -d

# AMD ROCm profile

docker compose -f deploy/docker-compose.yml --profile rocm up -d

Worker-only modes (worker-gpu, worker-rocm) omit the HTTP port and run headless GPU workers intended to join a remote control plane. These require the OMNIVOICE_WORKER_TOKEN environment variable as specified in the Compose file.

ARM64 and Apple Silicon Deployment

The official images target linux/amd64 only. On ARM64 hosts (including Apple Silicon), force AMD64 emulation:

export DOCKER_DEFAULT_PLATFORM=linux/amd64
docker compose -f deploy/docker-compose.yml --profile cpu pull
docker compose -f deploy/docker-compose.yml --profile cpu up -d

GPU acceleration is unavailable under emulation; use the CPU profile only.

Essential Runtime Environment Variables

Configure these variables when you deploy the VoiceStudio backend to control authentication, networking, and data paths:

Variable Purpose Default/Required
OMNIVOICE_API_KEY Administrator API key for privileged endpoints (/system/*, /api/settings/*) Required
OMNIVOICE_SERVER_MODE Set to 1 to disable the desktop-only loopback origin gate for headless deployment 1 (set in image)
OMNIVOICE_BIND_HOST Interface uvicorn binds to inside the container; must be 0.0.0.0 for external access 0.0.0.0
OMNIVOICE_PUBLIC_API_BASE Overrides the API base URL when behind a reverse proxy Optional

Other paths (database location, model cache) derive from mounted volumes and defaults defined in backend/core/config.py. The admin key validation is implemented in backend/core/auth.py.

Data Persistence Architecture

The container stores state in two Docker volumes:

  • omnivoice-data – SQLite database, user voices, and application data mapped to /app/omnivoice_data
  • ~/.cache/huggingface – HuggingFace model cache mapped to /root/.cache/huggingface

These volumes ensure that downloaded models and user data survive container restarts and updates.

Troubleshooting Deployment Issues

Loopback Origin Errors

If you encounter origin errors accessing the UI, verify OMNIVOICE_SERVER_MODE=1 is set. This relaxes the desktop-only loopback gate implemented in backend/main.py. When fronting the container with a custom reverse proxy, set OMNIVOICE_SERVER_MODE=0 to re-enable strict origin checking.

GPU Detection Failures

Verify CUDA or ROCm availability inside the container:

docker exec <container> python3 -c "import torch; print(torch.cuda.is_available())"

For NVIDIA, ensure you passed --gpus all. For AMD ROCm, confirm the device nodes (/dev/kfd, /dev/dri) are mounted and you used the :rocm image tag.

LAN Access Configuration

By default, the container binds to 127.0.0.1:3900 for security. To expose the VoiceStudio backend to your local network, modify the port mapping in deploy/docker-compose.yml from 127.0.0.1:3900:3900 to 0.0.0.0:3900:3900, or use -p 0.0.0.0:3900:3900 in your docker run command.

For additional troubleshooting, refer to the official documentation at /docs/install/docker.md in the repository.

Summary

  • Container image: Available at ghcr.io/debpalash/omnivoice-studio:latest (NVIDIA/CPU) and :rocm (AMD)
  • Required volumes: omnivoice-data for application state and ~/.cache/huggingface for models
  • Authentication: Generate OMNIVOICE_API_KEY with secrets.token_urlsafe(32) before first run
  • Networking: Exposes port 3900; set OMNIVOICE_SERVER_MODE=1 for headless deployment
  • Entrypoint: uvicorn backend.main:app serves the FastAPI backend and compiled React frontend
  • Source files: Configuration logic resides in backend/core/config.py and authentication in backend/core/auth.py

Frequently Asked Questions

What hardware acceleration does VoiceStudio support?

VoiceStudio supports NVIDIA GPUs via CUDA (using the default image with --gpus all), AMD GPUs via ROCm (using the :rocm tag with /dev/kfd and /dev/dri devices), and CPU-only inference. The deploy/docker-compose.yml provides specific profiles for each hardware configuration. GPU acceleration is not available on ARM64 hosts running under emulation.

How do I secure the VoiceStudio backend deployment?

Security requires setting the OMNIVOICE_API_KEY environment variable to a cryptographically secure random string generated with Python's secrets.token_urlsafe(32). This key grants administrative access to system settings and diagnostics endpoints as enforced by the validation logic in backend/core/auth.py. Always generate this key before starting the container and store it securely.

Why is the VoiceStudio UI inaccessible from another machine?

The default configuration binds to 127.0.0.1:3900 for security. To access the VoiceStudio backend from other machines on your network, change the port mapping to 0.0.0.0:3900:3900 in your docker run command or in deploy/docker-compose.yml. Ensure OMNIVOICE_SERVER_MODE=1 is set to disable the loopback-origin gate that restricts access to localhost only.

Can I run VoiceStudio on Apple Silicon Macs?

Yes, but only using CPU inference via AMD64 emulation. Set DOCKER_DEFAULT_PLATFORM=linux/amd64 before pulling the image, and use the CPU profile in Docker Compose. The container will run under emulation with reduced performance, and GPU acceleration is unavailable on Apple Silicon because the images do not include ARM64 builds or Metal support.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →