How to Set Up the VoiceStudio Backend Locally: FastAPI Setup Guide
To set up the VoiceStudio backend locally, clone the repository, install Python 3.11+ dependencies using uv sync, optionally create a .env file with OMNIVOICE_API_KEY for admin access, and launch the server with uv run python backend/main.py on port 3900.
VoiceStudio is an open-source voice synthesis and recognition platform built with FastAPI. This guide walks you through exactly how to set up the VoiceStudio backend locally using the source code from the debpalash/VoiceStudio repository.
Prerequisites
Before starting, ensure your environment meets the following requirements:
- Python 3.11 or higher installed on your system
- uv installed as the Python package manager (a fast, modern replacement for
pipandvenv) - Git for cloning the repository
Step-by-Step Local Installation
Follow these sequential steps to get the VoiceStudio backend running on localhost:3900.
Clone the Repository
First, download the source code from GitHub and navigate into the project directory:
git clone https://github.com/debpalash/VoiceStudio.git
cd VoiceStudio
Install Dependencies with uv
VoiceStudio uses uv for reproducible environment management. Running the sync command creates a virtual environment and installs all required packages:
uv sync
This command reads the project dependencies and constructs a .venv directory with the exact package versions specified in the repository's lock file.
Configure Environment Variables (Optional)
For local development, most features work without configuration. However, if you need access to admin-only routes, you must set the OMNIVOICE_API_KEY environment variable.
Create a .env file in the project root:
cat <<EOF > .env
OMNIVOICE_API_KEY=$(python -c "import secrets, sys; print(secrets.token_urlsafe(32))")
EOF
The backend/main.py file automatically loads variables from .env on startup. The OMNIVOICE_SERVER_MODE flag is used for Docker deployments and is not required for local runs.
Start the FastAPI Server
The entry point is located at backend/main.py. You have two methods to launch the server:
Method 1: Using uvicorn directly
uv run python -m uvicorn backend.main:app --host 127.0.0.1 --port 3900
Method 2: Using the bundled convenience script
uv run python backend/main.py
When the server starts, the lifespan context manager (lines 31-35 in backend/main.py) initializes the application. This context performs admin-key validation via core.auth.validate_server_admin_key and triggers deferred startup sequences in backend/services/model_manager.py to pre-load TTS and ASR engines.
Verify the Installation
Confirm the backend is responding correctly by querying the voices endpoint:
curl http://127.0.0.1:3900/v1/audio/voices
A successful response returns a JSON object containing available voice models:
{
"voices": [...]
}
Key Backend Components
Understanding these core files helps with troubleshooting and customization:
backend/main.py– The primary entry point that configures import paths, loads environment variables, registers API routers frombackend/api/, and constructs the FastAPI application.backend/api/– Directory containing REST route handlers for audio processing, TTS (text-to-speech), ASR (automatic speech recognition), and OpenAI-compatible endpoints.backend/services/model_manager.py– Handles heavy initialization tasks including model loading, engine warm-up, and watermark pool startup that runs after the server binds to the port.backend/core/config.py– Central configuration hub defining log paths, cache directories, and default port settings (3900).
Docker Alternative
If you prefer containerized deployment instead of a local Python environment, the repository includes Docker configuration in deploy/docker-compose.yml. The Docker setup requires the same OMNIVOICE_API_KEY environment variable for administrative functions. See docs/install/docker.md for container-specific instructions.
Summary
- VoiceStudio requires Python 3.11+ and uses uv for dependency management via
uv sync. - The backend entry point is
backend/main.py, which launches a FastAPI service on port 3900. - Set
OMNIVOICE_API_KEYin a.envfile only if you need admin route access; it is not required for basic local operation. - Start the server using
uv run python backend/main.pyor the uvicorn equivalent. - Verify functionality by requesting
/v1/audio/voicesvia curl or a FastAPI TestClient.
Frequently Asked Questions
What is the minimum Python version required for VoiceStudio?
VoiceStudio requires Python 3.11 or higher. This requirement is enforced by the project's dependency specification and the modern Python features used in the FastAPI implementation.
Why do I need an API key for local development?
You only need the OMNIVOICE_API_KEY if you intend to use admin-only routes or the administrative UI. For standard TTS, ASR, and audio processing endpoints, the backend runs without any authentication when executed locally. The key is validated through the validate_server_admin_key function in the lifespan context manager.
Can I use pip instead of uv to install dependencies?
While the repository is optimized for uv (as evidenced by the uv sync command and project structure), you can theoretically install requirements manually if you extract them from the project's configuration files. However, using uv ensures you match the exact environment specified by the maintainers and avoids dependency conflicts.
How do I change the default port from 3900?
The default port is defined in backend/core/config.py. You can override it by setting the appropriate environment variable in your .env file or by passing the --port argument when launching with uvicorn (e.g., uvicorn backend.main:app --port 8080).
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →