Voice-Pro Python Dependencies: The Complete AI Dubbing Stack Explained

Voice-Pro relies on 30+ specialized Python libraries including Gradio 6.20.0 for the UI, Faster-Whisper and OpenAI Whisper for ASR, F5-TTS and CosyVoice for voice cloning, and PyTorch 2.8.0 as its deep learning runtime, all declared in pyproject.toml with optional GPU/CPU extras.

Voice-Pro is an open-source AI dubbing application built exclusively for Python 3.12 that orchestrates media processing, speech recognition, translation, and voice synthesis through a Gradio web interface. Understanding the main Python dependencies for Voice-Pro is essential for developers contributing to the codebase or deploying custom instances. The project's dependency graph is strictly managed in the repository's pyproject.toml file, which defines everything from audio I/O utilities to neural TTS engines.

Core Dependency Architecture

The Voice-Pro Python dependencies are organized into logical layers that mirror the application's data pipeline: media ingestion, speech recognition, translation, and voice synthesis.

Media Acquisition and Processing

At the foundation, Voice-Pro handles video downloads and audio manipulation through industrial-strength libraries. yt-dlp (line 30) downloads source media from streaming platforms, while ffmpeg-python (line 29) provides bindings for codec operations. For speaker separation, the project uses demucs (line 31), and pydub (line 32) handles audio segment operations. Audio analysis relies on librosa (line 36) with soundfile (line 34) and soundcard (line 35) managing I/O operations.

Automatic Speech Recognition (ASR)

Voice-Pro implements a multi-engine ASR stack supporting various Whisper variants. The core packages include openai-whisper==20250625 (line 25) for baseline transcription, faster-whisper==1.2.1 (line 26) for optimized CPU/GPU inference, and whisper-timestamped==1.15.9 (line 27) for precise word-level alignment. In app/abus_asr_faster_whisper.py, these engines are wrapped to provide a unified interface for the dubbing pipeline.

Translation and Subtitle Generation

For multilingual support, the application combines free and enterprise translation services. edge-tts>=7.2.8 (line 45) provides Microsoft's neural voices, while deep-translator (line 47) offers aggregation across multiple free translation APIs. Optional Azure integration uses azure-ai-translation-text==1.0.0b1 (line 50). Subtitle manipulation relies on pysubs2 (line 48), with lingua-language-detector (line 49) handling automatic language identification.

Text-to-Speech and Voice Cloning

The TTS layer represents Voice-Pro's most complex dependency set, supporting multiple neural architectures. f5-tts==1.1.21 (line 59) provides high-quality voice cloning, while the CosyVoice integration (managed in app/abus_tts_cosyvoice.py) requires torchcodec>=0.7,<0.8 (line 61), conformer==0.3.2, diffusers==0.29.0, and transformers==5.13.0 (lines 63-74). Additional voice engines include kokoro==0.9.4, misaki[en,ja,zh]==0.9.4 for Japanese/Chinese support, and pyopenjtalk-plus (line 81) for Japanese phonemization. Text normalization uses wetext>=0.1.4 (line 74).

Web Interface and System Utilities

The Gradio-based UI depends on gradio==6.20.0 (line 39) locked to a specific version for stability, with numpy>=2.1,<2.5 (line 40) handling numerical arrays. System introspection uses py-cpuinfo (line 12), while structlog (line 11) provides structured logging throughout the application. Configuration management relies on python-dotenv (line 21), and the console interface uses rich (line 19) with markdown (line 18) for formatted output.

Deep Learning Runtime

Voice-Pro supports both CPU and GPU execution through optional dependency extras defined in pyproject.toml lines 92-103. The GPU extra installs torch==2.8.0, torchvision==0.23.0, torchaudio==2.8.0, and onnxruntime-gpu==1.26.0. The CPU extra substitutes onnxruntime==1.26.0 while keeping identical PyTorch versions. This design allows the uv package manager to resolve the appropriate runtime based on deployment target.

Implementation Examples

Below are practical implementations demonstrating how Voice-Pro utilizes its core dependencies.

Structured Logging with structlog

Voice-Pro uses structured logging for observability across the dubbing pipeline:

import structlog

logger = structlog.get_logger(__name__)

def process_audio(file_path: str):
    logger.info("audio_processing_started", path=file_path, stage="asr")
    # Processing logic here

    logger.info("audio_processing_completed", path=file_path)

Media Download with yt-dlp

The application downloads source media using yt-dlp as configured in the workspace:

from yt_dlp import YoutubeDL

def download_media(url: str, output_dir: str = "workspace"):
    ydl_opts = {
        "outtmpl": f"{output_dir}/%(title)s.%(ext)s",
        "format": "bestaudio/best"
    }
    with YoutubeDL(ydl_opts) as ydl:
        ydl.download([url])

Fast Transcription with Faster-Whisper

The ASR implementation in app/abus_asr_faster_whisper.py leverages the optimized Whisper backend:

from faster_whisper import WhisperModel

def transcribe_audio(audio_path: str):
    # Automatically uses CUDA when GPU extras are installed

    model = WhisperModel("large-v2", device="auto", compute_type="float16")
    segments, info = model.transcribe(audio_path, beam_size=5)
    
    return {
        "language": info.language,
        "text": " ".join([segment.text for segment in segments])
    }

Translation Pipeline

Voice-Pro supports multiple translation backends through deep-translator:

from deep_translator import GoogleTranslator

def translate_text(text: str, source: str = "auto", target: str = "ko"):
    translator = GoogleTranslator(source=source, target=target)
    return translator.translate(text)

Gradio Interface Components

The UI assembly in app/abus_app_voice.py constructs the dubbing workflow interface:

import gradio as gr

def create_dubbing_interface():
    with gr.Blocks(title="Voice-Pro") as demo:
        with gr.Row():
            input_audio = gr.Audio(label="Source Audio")
            output_audio = gr.Audio(label="Dubbed Output")
        
        process_btn = gr.Button("Generate Dubbing")
        process_btn.click(
            fn=dubbing_pipeline,
            inputs=input_audio,
            outputs=output_audio
        )
    
    return demo

Summary

  • Voice-Pro's dependency stack is defined exclusively in pyproject.toml, managed by the uv package manager for reproducible builds.
  • The ASR layer combines Faster-Whisper, OpenAI Whisper, and Whisper-Timestamped for accurate speech-to-text with alignment.
  • Neural voice synthesis relies on F5-TTS, CosyVoice, and Kokoro, supported by PyTorch 2.8.0 and Transformers 5.13.0.
  • Media processing requires yt-dlp, ffmpeg-python, demucs, and librosa for robust audio pipeline handling.
  • The web interface locks Gradio 6.20.0 with NumPy 2.x for consistent UI behavior across platforms.
  • Optional extras in pyproject.toml lines 92-103 allow selection between ONNX Runtime GPU or CPU variants without changing application code.

Frequently Asked Questions

What Python version does Voice-Pro require?

Voice-Pro requires Python 3.12 exclusively. The pyproject.toml specifies this constraint strictly, and attempting installation on earlier versions will fail during dependency resolution.

How does Voice-Pro handle GPU acceleration?

GPU support is implemented through optional extras defined in pyproject.toml lines 92-103. Installing with the gpu extra includes onnxruntime-gpu==1.26.0 and CUDA-enabled PyTorch wheels, while the default CPU extra uses standard onnxruntime. The application auto-detects available hardware in app/gradio_gulliver.py and adjusts compute types accordingly.

Where are the ASR and TTS models configured?

Speech recognition models are configured in app/abus_asr_faster_whisper.py, which wraps the Faster-Whisper library. TTS engines are managed in app/abus_tts_cosyvoice.py for CosyVoice integration and related modules for F5-TTS. Both import their underlying libraries (e.g., faster_whisper, f5_tts) directly from the dependencies declared in pyproject.toml lines 25-27 and 59.

Can I install Voice-Pro dependencies without the GPU libraries?

Yes. The repository uses optional dependency groups. Installing without the gpu extra will pull torch, torchvision, and torchaudio CPU wheels plus onnxruntime instead of onnxruntime-gpu, significantly reducing installation size while maintaining full functionality on CPU-only systems.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →