Common Troubleshooting Steps for Voice-Pro Issues: A Complete Guide
Delete the installer_files/ directory and rerun the launcher to fix most startup crashes, or adjust compute types and model sizes to resolve CUDA memory errors.
Voice-Pro is an open-source, three-layer Gradio web application that orchestrates video downloading, ASR (Automatic Speech Recognition), translation, and TTS (Text-to-Speech) pipelines. Because it integrates external tools like ffmpeg, CUDA drivers, and Azure cloud services, issues often manifest as UI error toasts or console tracebacks. This guide covers the most frequent Voice-Pro problems and their fixes based on the actual source code in abus-aikorea/voice-pro.
Understanding the Voice-Pro Architecture
Voice-Pro follows a modular Gradio architecture defined in app/abus_app_voice.py. The entry point start-voice.py loads user configuration from src/config.py and assembles the UI from tab modules (app/tab_*.py), which delegate processing to core functions (app/abus_*.py). Errors typically bubble up from the processing layer through Gradio controllers (app/gradio_*.py) and are logged via structlog.
Knowing this flow helps isolate whether a failure stems from environment setup (installer/virtual-env), configuration (.env files), or runtime resource limits (GPU memory).
Fixing Installation and Startup Failures
Red Error Toasts or Crashes on Launch
If Voice-Pro displays a red error toast immediately or crashes during startup, the cause is usually corrupted or missing files in the installer_files/ directory. This folder contains the uv package manager, Python interpreter, and installed dependencies.
Recommended fix: Delete the entire installer_files/ directory and rerun start.bat (Windows) or start.sh (Linux/macOS). The launcher will redownload uv, create a fresh virtual environment, and reinstall all dependencies.
# Force a clean reinstall programmatically
import shutil
import os
import subprocess
installer_dir = os.path.join(os.getcwd(), "installer_files")
if os.path.isdir(installer_dir):
shutil.rmtree(installer_dir) # Remove corrupted environment
subprocess.run(["start.bat"], shell=True) # Re-run launcher
Browser Does Not Open Automatically
When the command window finishes but the browser fails to launch, the auto-launch flag may be disabled, or the Windows command window may have closed prematurely.
Recommended fix: Manually navigate to http://127.0.0.1:7870 in your browser, or close the terminal and rerun start.bat to trigger the launch sequence again.
Resolving CUDA and GPU Memory Errors
CUDA Out-of-Memory Errors
Selecting a large Whisper model, enabling denoise level 2, or using float compute precision can exhaust GPU RAM, causing the ASR pipeline to fail.
Recommended fix: Reduce the denoise level to 0 or 1, or switch the compute type to int (quantized) in the Gradio UI. The backend applies this setting in modules like app/abus_asr_faster_whisper.py.
# Example: Configuring compute type to reduce VRAM usage
# In the UI, set "Compute Type" dropdown to "int" for quantized inference
# This is read by the ASR backend during model initialization
GPU Not Detected
Voice-Pro relies on the GPU_CHOICE environment variable or a gpu_choice.txt file in installer_files/ to determine whether to use CUDA or CPU.
Recommended fix: Delete installer_files/gpu_choice.txt to force re-detection, or explicitly set the environment variable before running the start script:
GPU_CHOICE=Gfor GPU (CUDA)GPU_CHOICE=Cfor CPU-only mode
Improving Subtitle and Audio Quality
Poor Subtitle Accuracy
Small Whisper models (e.g., tiny or base) or unsuitable compute types produce low-quality transcriptions.
Recommended fix: Switch to a larger model such as large-v3-turbo and ensure the compute type is set to float for higher fidelity rather than quantized int modes. This trades processing speed for accuracy as documented in the README troubleshooting section.
Fixing Translation and TTS Service Failures
Intermittent Translation or TTS Errors
The default Google-based translation endpoints are subject to rate limiting and geographic blocking, causing sporadic failures in the dubbing pipeline.
Recommended fix: Configure Azure Translator and Azure Speech Services by copying .env.example to .env and filling in your API keys. The application validates these credentials in app/abus_config.py using helper functions like get_azure_speech_key() and get_azure_translator_key().
# Example: Loading Azure credentials from .env
from app.abus_config import get_azure_speech_key, get_azure_translator_key
try:
speech_key = get_azure_speech_key() # Raises ValueError if missing
translator_key = get_azure_translator_key() # Raises ValueError if missing
print("Azure credentials loaded successfully")
except ValueError as e:
print(f"Configuration error: {e}")
The UI assembly logic in app/abus_app_voice.py calls azure_text_api_working() from app/abus_genuine.py to determine whether to display the Azure TTS tab or fall back to limited offline options.
Addressing FFmpeg and Model Download Issues
FFmpeg-Related Errors
Missing codecs or ffmpeg not being on the system PATH causes video processing failures.
Recommended fix: The start.bat script automatically downloads a portable ffmpeg binary into installer_files/ffmpeg/. If you maintain a custom ffmpeg installation, ensure the binary is reachable from your system PATH environment variable.
Stalled or Corrupted Model Downloads
Network interruptions during initial setup can leave partial model files in the model/ directory.
Recommended fix: Re-run the start script. The downloader will resume or redownload missing model files. All filesystem paths—including model/, workspace/, and installer_files/—are resolved through app/abus_path.py to handle platform-specific directory conventions.
Where Fixes Live in the Codebase
Understanding the specific files responsible for error handling helps with advanced debugging:
app/abus_app_voice.py– Top-level UI assembly that wires tabs and checks Azure service availability viaazure_text_api_working().app/abus_config.py– Reads.envvariables viapython-dotenvand raises clear exceptions if Azure keys are malformed or missing.app/abus_path.py– Centralized path resolution forinstaller_files/,model/, andworkspace/directories.app/abus_genuine.py– Detects whether Azure services are properly configured and enabled.app/gradio_gulliver.py– Core "Dubbing Studio" pipeline controller where processing errors are logged.start-voice.py– Entry point that handlesGPU_CHOICElogic and launches the Gradio interface.
Summary
- Clean reinstalls fix most startup crashes by removing corrupted
installer_files/directories. - CUDA OOM errors are resolved by reducing denoise levels or switching to
intcompute types. - Azure credentials must be configured in
.env(copied from.env.example) to avoid translation/TTS rate limits. - GPU detection relies on the
GPU_CHOICEenvironment variable or deleting stalegpu_choice.txtfiles. - Model and ffmpeg paths are managed automatically by the launcher, but custom installations require proper
PATHconfiguration.
Frequently Asked Questions
How do I completely reset Voice-Pro to factory settings?
Delete the installer_files/ directory in the project root and rerun start.bat or start.sh. This forces the launcher to redownload uv, recreate the Python virtual environment, and reinstall all dependencies without touching your workspace files or downloaded models.
Why does Voice-Pro fail with a CUDA out-of-memory error on large files?
The selected Whisper model or denoise level 2 processing exceeds your GPU's VRAM. Open the Gradio UI settings and reduce the denoise level to 0 or 1, or change the compute type from float to int (quantized) to reduce memory consumption by approximately 50%.
How do I enable Azure Text-to-Speech to avoid Google rate limits?
Copy .env.example to .env in the project root, then add your Azure Speech key and Translator key to the file. The app/abus_config.py module will automatically load these values on startup, and app/abus_app_voice.py will enable the Azure TTS tab if the credentials are valid.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →