How to Use Environment Variables to Tune VoiceStudio Performance: Complete Configuration Guide
VoiceStudio exposes ten well-documented environment variables that control batch processing width, timeout budgets, external tool paths, and text normalization—allowing you to optimize throughput, latency, and resource consumption without modifying source code.
VoiceStudio's runtime behavior is governed by environment variables read at startup from core/user_env.py and propagated throughout the FastAPI application. By tuning these variables, you can adapt the system to hardware ranging from single-GPU laptops to multi-GPU server farms. This guide covers each variable's purpose, default values, and where they are applied in the codebase.
Core Performance Tuning Variables
Batch Processing Width
The OMNIVOICE_DUB_BATCH_WIDTH variable overrides the auto-detected batch width for TTS generation. A larger value amortizes encoder/decoder setup costs across more segments, significantly speeding up multi-segment dubbing when VRAM permits.
- Default: Auto-detected from host VRAM, capped at 16
- Valid range: 1–16 (values outside this range are clamped)
- Applied in:
backend/api/routers/batch.pyat lines 46–53
When unset or invalid, _native_batch_width() derives the width from available GPU memory. For CPU-only deployments, the system falls back to a conservative default.
Generation Timeouts
Three interrelated variables control how long the system waits for TTS output:
| Variable | Purpose | Default |
|---|---|---|
OMNIVOICE_GENERATE_TIMEOUT_S |
Global floor for GPU generation | 300 seconds |
OMNIVOICE_CPU_GENERATE_TIMEOUT_S |
Separate floor for CPU paths | 600 seconds |
OMNIVOICE_GGUF_GENERATE_TIMEOUT_S |
Fine-grained control for GGUF engines | Falls back to generic timeout |
These values form a base budget per segment. The actual timeout adds per-segment length overage, ensuring longer texts receive proportionally more time. All three are consumed by generate_timeout_s() in services/model_manager.py, which feeds the GPU-pool guard.
Model Loading Timeout
The OMNIVOICE_MODEL_LOAD_TIMEOUT prevents a hung model load from indefinitely blocking the worker pool. Its default is derived from the generation timeout plus a 100-second safety margin.
- Location:
services/model_manager.py - Behavior: Wraps import/initialization; raises
TimeoutErroron breach
Operational Reliability Variables
External Tool Paths
When ffmpeg or ffprobe are not on $PATH or you require a specific build, use:
FFMPEG_PATHFFPROBE_PATH
These are read by find_ffmpeg() in backend/services/media_tools.py (lines 98–103). Explicit paths eliminate costly PATH lookups on constrained systems and enable reproducible deployments with pinned binary versions.
Text Normalization Toggle
The OMNIVOICE_TEXT_NORMALIZATION flag controls whether normalize_for_tts() runs before generation. Disabling it removes regex-heavy preprocessing, reducing CPU load for large batches at potential cost to output quality.
- Default:
"1"(enabled) - Read in:
backend/services/text_normalization.pyat line 68
IndexTTS Socket Timeout
For deployments using the IndexTTS side-car service, OMNIVOICE_INDEXTTS_RECV_TIMEOUT_S sets the socket receive timeout. This prevents a stalled subprocess from hanging the entire pipeline.
- Default: 60 seconds
- Consumed in:
backend/services/subprocess_backend.py
Complete Configuration Examples
Production Deployment (.env file)
Create a .env file alongside the VoiceStudio executable:
# High-throughput GPU server configuration
OMNIVOICE_DUB_BATCH_WIDTH=16 # Maximum native batch width
OMNIVOICE_GENERATE_TIMEOUT_S=400 # Extended GPU budget for long content
OMNIVOICE_MODEL_LOAD_TIMEOUT=600 # 10 minutes for large model loads
FFMPEG_PATH=/opt/ffmpeg/6.0/bin/ffmpeg # Pinned ffmpeg version
Development/Low-Resource Setup
# CPU-only laptop configuration
OMNIVOICE_DUB_BATCH_WIDTH=2 # Conservative batching
OMNIVOICE_CPU_GENERATE_TIMEOUT_S=900 # Generous timeout for slow CPU
OMNIVOICE_TEXT_NORMALIZATION=0 # Faster preprocessing
Runtime Override (Shell)
export OMNIVOICE_DUB_BATCH_WIDTH=8
export OMNIVOICE_GENERATE_TIMEOUT_S=500
voice-studio # launches with custom limits
Programmatic Access Pattern
The internal logic for reading these variables follows this pattern (mirrored from backend/api/routers/batch.py):
import os
# Batch width with clamping to valid range
raw_width = int(os.getenv("OMNIVOICE_DUB_BATCH_WIDTH", "0") or 0)
batch_width = (
max(1, min(16, raw_width))
if raw_width else
None # Triggers auto-detect from VRAM
)
# Timeout floors
generate_timeout = float(
os.getenv("OMNIVOICE_GENERATE_TIMEOUT_S", "300")
)
cpu_generate_timeout = float(
os.getenv("OMNIVOICE_CPU_GENERATE_TIMEOUT_S", "600")
)
# Custom tool paths
ffmpeg_path = os.getenv("FFMPEG_PATH") # None if unset
Temporary Override for Testing
import os
from fastapi.testclient import TestClient
from main import app
client = TestClient(app)
# Isolate environment change to single test
original = os.environ.get("OMNIVOICE_DUB_BATCH_WIDTH")
try:
os.environ["OMNIVOICE_DUB_BATCH_WIDTH"] = "12"
response = client.post(
"/batch/enqueue",
files={"video": ("sample.mp4", open("sample.mp4", "rb"))},
data={"langs": "en,es"},
)
assert response.status_code == 200
finally:
if original is None:
os.environ.pop("OMNIVOICE_DUB_BATCH_WIDTH", None)
else:
os.environ["OMNIVOICE_DUB_BATCH_WIDTH"] = original
Architectural Flow: How Variables Propagate
- Startup:
core/user_env.pyloads.envintoos.environbefore the FastAPI application initializes - Batch enqueue:
POST /batch/enqueuetriggers_native_batch_width(), which checksOMNIVOICE_DUB_BATCH_WIDTHbefore VRAM auto-detection - Timeout calculation: Each segment's budget is computed by
generate_timeout_s()using the environment floor plus length-based overage - Model loading: First-use model initialization respects
OMNIVOICE_MODEL_LOAD_TIMEOUTvia a guarded wrapper - Media processing:
find_ffmpeg()resolves binaries usingFFMPEG_PATH/FFPROBE_PATHwhen present - Text pipeline:
normalize_for_tts()is conditionally executed based onOMNIVOICE_TEXT_NORMALIZATION
Key Source Files Reference
| File | Responsibility | Link |
|---|---|---|
backend/api/routers/batch.py |
Batch width override logic and capping | batch.py#L46-L53 |
services/model_manager.py |
All timeout calculations and model-load guarding | Search: generate_timeout_s |
backend/services/text_normalization.py |
Normalization toggle check | text_normalization.py#L68 |
backend/services/media_tools.py |
External tool path resolution | media_tools.py#L98-L103 |
core/user_env.py |
Environment loading bootstrap | Search: user_env |
tests/test_generate_timeout_settings_1787.py |
Validation of timeout behavior | test file |
tests/test_batch_width_and_budget.py |
Batch width fallback validation | test file |
Summary
OMNIVOICE_DUB_BATCH_WIDTHcontrols TTS parallelism (1–16, auto-detected if unset)- Timeout variables (
GENERATE_TIMEOUT_S,CPU_GENERATE_TIMEOUT_S,GGUF_GENERATE_TIMEOUT_S,MODEL_LOAD_TIMEOUT) define per-operation budgets and protect pool health FFMPEG_PATHandFFPROBE_PATHenable custom external tool binariesOMNIVOICE_TEXT_NORMALIZATIONtrades preprocessing quality for CPU efficiencyOMNIVOICE_INDEXTTS_RECV_TIMEOUT_Sisolates side-car service failures
Adjust these VoiceStudio environment variables to match your hardware capabilities, optimize for latency or throughput, and maintain stability under load.
Frequently Asked Questions
What happens if I set OMNIVOICE_DUB_BATCH_WIDTH higher than 16?
Values above 16 are clamped to 16 by _native_batch_width() in backend/api/routers/batch.py. The function enforces min(value, _MAX_BATCH_WIDTH) where _MAX_BATCH_WIDTH = 16. Invalid values (non-integers, negative numbers) fall back to auto-detection from VRAM.
How do CPU and GPU timeouts interact?
OMNIVOICE_CPU_GENERATE_TIMEOUT_S is consulted only when the generation path detects CPU-only execution. For GPU paths, OMNIVOICE_GENERATE_TIMEOUT_S applies. If neither is set, the defaults are 600 seconds for CPU and 300 seconds for GPU. The GGUF-specific timeout overrides both when using GGUF-based engines.
Can I disable environment variable loading from .env entirely?
VoiceStudio loads core/user_env.py implicitly from main.py, which uses standard python-dotenv behavior. To bypass .env loading, ensure no .env file exists in the working directory or set variables directly in the process environment—these take precedence over file-based values.
Where should I report unexpected behavior with environment variables?
The test files test_generate_timeout_settings_1787.py and test_batch_width_and_budget.py define the expected behavior and safety floors. If you observe deviations, compare your configuration against these specifications and check the relevant source file links provided above before filing an issue.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →