Core Features of Voice-Pro: A Complete AI Dubbing and Voice Generation Platform
Voice-Pro is an open-source Gradio web application that delivers end-to-end AI-powered dubbing, integrating YouTube downloading, source separation, speech recognition, translation, and voice cloning through a modular three-layer architecture.
The core features of Voice-Pro provide a comprehensive multimedia pipeline for creators and developers. Developed by abus-aikorea, this platform combines production-grade AI models—including Whisper, Demucs, and multiple TTS engines—within a unified interface. Whether generating subtitles for 90+ languages or cloning voices with zero-shot learning, Voice-Pro separates UI components from processing logic to ensure scalability and ease of use.
End-to-End Dubbing Studio
The Dubbing Studio represents the primary workflow of Voice-Pro, handling complete media transformation from ingestion to final output. This feature manages YouTube video downloading via abus_downloader.py, vocal separation using Demucs algorithms in abus_demucs.py, and speech-to-text conversion through Faster-Whisper implementations.
According to the source code, the UI layer resides in app/tab_gulliver.py, while the orchestration logic connecting ASR, translation, and TTS modules lives in app/abus_app_voice.py. Users can process entire videos without leaving the interface, as the system automatically coordinates between abus_asr_faster_whisper.py for transcription and the translation wrappers for multilingual output.
Speech Recognition with Whisper Integration
Voice-Pro provides advanced ASR (Automatic Speech Recognition) capabilities supporting 90+ languages with word-level timestamp precision. The platform offers multiple engine options through app/abus_asr_faster_whisper.py and app/abus_asr_whisper_timestamped.py, allowing users to select between standard Whisper, Faster-Whisper, or Whisper-Timestamped variants.
Key implementation details include optional audio denoising pre-processing and compute-type selection (CPU vs. CUDA) configured through the controller layer. These engines generate subtitle files with precise timing information essential for dubbing workflows.
Multilingual Translation Engine
The translation infrastructure supports 100+ languages through dual backend options. By default, app/abus_translate_deep.py utilizes Deep-Translator's Google endpoint with built-in retry and back-off logic for free usage. For enterprise requirements, app/abus_translate_azure.py provides optional Azure Translator integration configured via environment variables in .env.
This module operates independently of the UI, allowing programmatic access for batch processing workflows beyond the Gradio interface.
Multi-Engine Voice Generation
Voice-Pro distinguishes itself through support for multiple TTS (Text-to-Speech) backends, each optimized for different use cases:
- Edge-TTS: Free access to 100+ languages and 400+ voices via Microsoft's Edge browser service, implemented in
app/gradio_tts_edge.pyandapp/abus_tts_edge.py - Azure-TTS: Optional cloud-based synthesis with neural voices
- F5-TTS: High-quality open-source synthesis through
app/gradio_tts_f5.py - CosyVoice: Zero-shot voice cloning including Fun-CosyVoice3 support via
app/gradio_tts_cosyvoice.pyand core logic inapp/abus_tts_cosyvoice.py - Kokoro: Lightweight efficient synthesis through
app/gradio_tts_kokoro.py
All TTS controllers inherit from base classes that standardize voice cloning workflows, enabling consistent API usage across different underlying models.
Voice Separation Technology
Clean vocal extraction is essential for professional dubbing results. Voice-Pro implements source separation through two algorithms:
- Demucs: Facebook's state-of-the-art music source separation, wrapped in
app/abus_demucs.py - MDX-Net: Alternative separation architecture in
app/abus_mdx.py
These modules process audio before ASR ingestion, significantly improving transcription accuracy by removing background music and noise from video sources.
AI Karaoke
The AI Karaoke feature repurposes the core ASR and TTS pipeline for real-time lyric synchronization. Implemented in app/tab_karaoke.py, this module generates instrumental tracks with overlaid synthesized vocals, maintaining timing alignment with the original media.
Cross-Platform Support
Voice-Pro targets Windows as its primary platform while maintaining compatibility structures for Linux and macOS (currently unverified). The installation system uses uv for reproducible Python environments, with automated scripts (configure.bat, configure.sh) handling dependency resolution and optional portable FFmpeg downloads.
Configuration persists in app/config-user.json5, loaded via src.config.UserConfig, while internationalization support through src/i18n/i18n.py enables UI localization.
Modular Three-Layer Architecture
The core features of Voice-Pro rely on a strict separation of concerns:
Tab Layer (app/tab_*.py): Builds Gradio interfaces and wires widgets to callbacks. For example, tab_gulliver.py creates the Dubbing Studio UI and connects upload handlers to controller methods.
Gradio Controllers (app/gradio_*.py): Stateful classes like GradioGulliver and GradioMSVoice maintain UI state and expose methods such as gradio_upload_source, gradio_whisper, and gradio_translate. These controllers bridge the visual interface and processing engines.
Core Modules (app/abus_*.py): Pure Python engines with no Gradio dependencies. This layer includes abus_asr_faster_whisper.py for transcription, abus_tts_cosyvoice.py for synthesis, abus_translate_deep.py for localization, and abus_downloader.py for media ingestion. All modules share structured logging via structlog and respect user configuration stored in app/config-user.json5.
Getting Started with Voice-Pro
Developers can launch the complete interface or access individual engines programmatically.
Launch the full web UI:
from app.abus_app_voice import create_ui
from src.config import UserConfig
user_cfg = UserConfig() # loads default + user config file
create_ui(user_cfg) # opens Gradio at http://127.0.0.1:7870
Direct TTS synthesis without the UI:
from app.abat_tts_edge import AzureTTS, EdgeTTS
from app.abu_path import path_workspace
text = "Hello, world!"
engine = EdgeTTS() # falls back to free Edge service
audio_path = engine.synthesize(text, voice="en-US-AriaNeural")
print(f"Synthesized audio saved to {audio_path}")
Note: These snippets assume the repository's virtual environment (created by the uv installer) is active.
Summary
- Voice-Pro integrates Dubbing Studio, Whisper Subtitles, Translation, and Voice Generation into a single Gradio-based platform.
- The architecture separates concerns into Tab, Controller, and Core layers for maintainability.
- ASR capabilities support 90+ languages via
abus_asr_faster_whisper.pyandabus_asr_whisper_timestamped.py. - TTS options include Edge-TTS, Azure-TTS, F5-TTS, CosyVoice, and Kokoro, with zero-shot voice cloning available.
- Source separation uses Demucs and MDX-Net via
abus_demucs.pyandabus_mdx.py. - Cross-platform installers use
uvfor environment management, with configuration stored inabus_config.py.
Frequently Asked Questions
What languages does Voice-Pro support for transcription and translation?
Voice-Pro supports 90+ languages for speech recognition through Whisper-based engines in abus_asr_faster_whisper.py, and 100+ languages for translation via the Deep-Translator integration in abus_translate_deep.py. Language selection is configurable through the Gradio interface or programmatically via the UserConfig class.
How does Voice-Pro handle voice cloning?
The platform provides zero-shot voice cloning through the CosyVoice implementation in app/abus_tts_cosyvoice.py, which includes support for Fun-CosyVoice3 models. Users can clone voices from short audio samples without additional training, with the Gradio controller managing model loading and inference through app/gradio_tts_cosyvoice.py.
Is Voice-Pro compatible with operating systems other than Windows?
While Windows is the primary supported platform, the repository includes Linux and macOS installation scripts (configure.sh, start.sh). However, these platforms are currently marked as unverified in the documentation. The portable installer relies on uv for cross-platform Python environment management.
What is the difference between the Dubbing Studio and AI Karaoke features?
The Dubbing Studio (implemented in tab_gulliver.py and abus_app_voice.py) provides a complete workflow for video translation and voice replacement. AI Karaoke (in tab_karaoke.py) uses the same underlying ASR and TTS engines but specifically targets real-time lyric synchronization and vocal generation for music applications, creating synchronized instrumental and vocal tracks.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →