# Voice-Pro Python Dependencies: The Complete AI Dubbing Stack Explained

> Explore the core Python dependencies for Voice-Pro, including Gradio, Faster-Whisper, F5-TTS, CosyVoice, and PyTorch. Understand the AI dubbing stack's essential libraries.

- Repository: [ABUS/voice-pro](https://github.com/abus-aikorea/voice-pro)
- Tags: deep-dive
- Published: 2026-08-03

---

**Voice-Pro relies on 30+ specialized Python libraries including Gradio 6.20.0 for the UI, Faster-Whisper and OpenAI Whisper for ASR, F5-TTS and CosyVoice for voice cloning, and PyTorch 2.8.0 as its deep learning runtime, all declared in [`pyproject.toml`](https://github.com/abus-aikorea/voice-pro/blob/main/pyproject.toml) with optional GPU/CPU extras.**

Voice-Pro is an open-source AI dubbing application built exclusively for Python 3.12 that orchestrates media processing, speech recognition, translation, and voice synthesis through a Gradio web interface. Understanding the main Python dependencies for Voice-Pro is essential for developers contributing to the codebase or deploying custom instances. The project's dependency graph is strictly managed in the repository's [`pyproject.toml`](https://github.com/abus-aikorea/voice-pro/blob/main/pyproject.toml) file, which defines everything from audio I/O utilities to neural TTS engines.

## Core Dependency Architecture

The Voice-Pro Python dependencies are organized into logical layers that mirror the application's data pipeline: media ingestion, speech recognition, translation, and voice synthesis.

### Media Acquisition and Processing

At the foundation, Voice-Pro handles video downloads and audio manipulation through industrial-strength libraries. **yt-dlp** (line 30) downloads source media from streaming platforms, while **ffmpeg-python** (line 29) provides bindings for codec operations. For speaker separation, the project uses **demucs** (line 31), and **pydub** (line 32) handles audio segment operations. Audio analysis relies on **librosa** (line 36) with **soundfile** (line 34) and **soundcard** (line 35) managing I/O operations.

### Automatic Speech Recognition (ASR)

Voice-Pro implements a multi-engine ASR stack supporting various Whisper variants. The core packages include **openai-whisper==20250625** (line 25) for baseline transcription, **faster-whisper==1.2.1** (line 26) for optimized CPU/GPU inference, and **whisper-timestamped==1.15.9** (line 27) for precise word-level alignment. In [`app/abus_asr_faster_whisper.py`](https://github.com/abus-aikorea/voice-pro/blob/main/app/abus_asr_faster_whisper.py), these engines are wrapped to provide a unified interface for the dubbing pipeline.

### Translation and Subtitle Generation

For multilingual support, the application combines free and enterprise translation services. **edge-tts>=7.2.8** (line 45) provides Microsoft's neural voices, while **deep-translator** (line 47) offers aggregation across multiple free translation APIs. Optional Azure integration uses **azure-ai-translation-text==1.0.0b1** (line 50). Subtitle manipulation relies on **pysubs2** (line 48), with **lingua-language-detector** (line 49) handling automatic language identification.

### Text-to-Speech and Voice Cloning

The TTS layer represents Voice-Pro's most complex dependency set, supporting multiple neural architectures. **f5-tts==1.1.21** (line 59) provides high-quality voice cloning, while the CosyVoice integration (managed in [`app/abus_tts_cosyvoice.py`](https://github.com/abus-aikorea/voice-pro/blob/main/app/abus_tts_cosyvoice.py)) requires **torchcodec>=0.7,<0.8** (line 61), **conformer==0.3.2**, **diffusers==0.29.0**, and **transformers==5.13.0** (lines 63-74). Additional voice engines include **kokoro==0.9.4**, **misaki[en,ja,zh]==0.9.4** for Japanese/Chinese support, and **pyopenjtalk-plus** (line 81) for Japanese phonemization. Text normalization uses **wetext>=0.1.4** (line 74).

### Web Interface and System Utilities

The Gradio-based UI depends on **gradio==6.20.0** (line 39) locked to a specific version for stability, with **numpy>=2.1,<2.5** (line 40) handling numerical arrays. System introspection uses **py-cpuinfo** (line 12), while **structlog** (line 11) provides structured logging throughout the application. Configuration management relies on **python-dotenv** (line 21), and the console interface uses **rich** (line 19) with **markdown** (line 18) for formatted output.

### Deep Learning Runtime

Voice-Pro supports both CPU and GPU execution through optional dependency extras defined in [`pyproject.toml`](https://github.com/abus-aikorea/voice-pro/blob/main/pyproject.toml) lines 92-103. The **GPU** extra installs **torch==2.8.0**, **torchvision==0.23.0**, **torchaudio==2.8.0**, and **onnxruntime-gpu==1.26.0**. The **CPU** extra substitutes **onnxruntime==1.26.0** while keeping identical PyTorch versions. This design allows the `uv` package manager to resolve the appropriate runtime based on deployment target.

## Implementation Examples

Below are practical implementations demonstrating how Voice-Pro utilizes its core dependencies.

### Structured Logging with structlog

Voice-Pro uses structured logging for observability across the dubbing pipeline:

```python
import structlog

logger = structlog.get_logger(__name__)

def process_audio(file_path: str):
    logger.info("audio_processing_started", path=file_path, stage="asr")
    # Processing logic here

    logger.info("audio_processing_completed", path=file_path)

```

### Media Download with yt-dlp

The application downloads source media using yt-dlp as configured in the workspace:

```python
from yt_dlp import YoutubeDL

def download_media(url: str, output_dir: str = "workspace"):
    ydl_opts = {
        "outtmpl": f"{output_dir}/%(title)s.%(ext)s",
        "format": "bestaudio/best"
    }
    with YoutubeDL(ydl_opts) as ydl:
        ydl.download([url])

```

### Fast Transcription with Faster-Whisper

The ASR implementation in [`app/abus_asr_faster_whisper.py`](https://github.com/abus-aikorea/voice-pro/blob/main/app/abus_asr_faster_whisper.py) leverages the optimized Whisper backend:

```python
from faster_whisper import WhisperModel

def transcribe_audio(audio_path: str):
    # Automatically uses CUDA when GPU extras are installed

    model = WhisperModel("large-v2", device="auto", compute_type="float16")
    segments, info = model.transcribe(audio_path, beam_size=5)
    
    return {
        "language": info.language,
        "text": " ".join([segment.text for segment in segments])
    }

```

### Translation Pipeline

Voice-Pro supports multiple translation backends through deep-translator:

```python
from deep_translator import GoogleTranslator

def translate_text(text: str, source: str = "auto", target: str = "ko"):
    translator = GoogleTranslator(source=source, target=target)
    return translator.translate(text)

```

### Gradio Interface Components

The UI assembly in [`app/abus_app_voice.py`](https://github.com/abus-aikorea/voice-pro/blob/main/app/abus_app_voice.py) constructs the dubbing workflow interface:

```python
import gradio as gr

def create_dubbing_interface():
    with gr.Blocks(title="Voice-Pro") as demo:
        with gr.Row():
            input_audio = gr.Audio(label="Source Audio")
            output_audio = gr.Audio(label="Dubbed Output")
        
        process_btn = gr.Button("Generate Dubbing")
        process_btn.click(
            fn=dubbing_pipeline,
            inputs=input_audio,
            outputs=output_audio
        )
    
    return demo

```

## Summary

- Voice-Pro's dependency stack is defined exclusively in [`pyproject.toml`](https://github.com/abus-aikorea/voice-pro/blob/main/pyproject.toml), managed by the `uv` package manager for reproducible builds.
- The ASR layer combines **Faster-Whisper**, **OpenAI Whisper**, and **Whisper-Timestamped** for accurate speech-to-text with alignment.
- Neural voice synthesis relies on **F5-TTS**, **CosyVoice**, and **Kokoro**, supported by **PyTorch 2.8.0** and **Transformers 5.13.0**.
- Media processing requires **yt-dlp**, **ffmpeg-python**, **demucs**, and **librosa** for robust audio pipeline handling.
- The web interface locks **Gradio 6.20.0** with **NumPy 2.x** for consistent UI behavior across platforms.
- Optional extras in [`pyproject.toml`](https://github.com/abus-aikorea/voice-pro/blob/main/pyproject.toml) lines 92-103 allow selection between **ONNX Runtime GPU** or CPU variants without changing application code.

## Frequently Asked Questions

### What Python version does Voice-Pro require?

Voice-Pro requires **Python 3.12** exclusively. The [`pyproject.toml`](https://github.com/abus-aikorea/voice-pro/blob/main/pyproject.toml) specifies this constraint strictly, and attempting installation on earlier versions will fail during dependency resolution.

### How does Voice-Pro handle GPU acceleration?

GPU support is implemented through optional extras defined in [`pyproject.toml`](https://github.com/abus-aikorea/voice-pro/blob/main/pyproject.toml) lines 92-103. Installing with the `gpu` extra includes **onnxruntime-gpu==1.26.0** and CUDA-enabled PyTorch wheels, while the default CPU extra uses standard **onnxruntime**. The application auto-detects available hardware in [`app/gradio_gulliver.py`](https://github.com/abus-aikorea/voice-pro/blob/main/app/gradio_gulliver.py) and adjusts compute types accordingly.

### Where are the ASR and TTS models configured?

Speech recognition models are configured in [`app/abus_asr_faster_whisper.py`](https://github.com/abus-aikorea/voice-pro/blob/main/app/abus_asr_faster_whisper.py), which wraps the **Faster-Whisper** library. TTS engines are managed in [`app/abus_tts_cosyvoice.py`](https://github.com/abus-aikorea/voice-pro/blob/main/app/abus_tts_cosyvoice.py) for CosyVoice integration and related modules for F5-TTS. Both import their underlying libraries (e.g., `faster_whisper`, `f5_tts`) directly from the dependencies declared in [`pyproject.toml`](https://github.com/abus-aikorea/voice-pro/blob/main/pyproject.toml) lines 25-27 and 59.

### Can I install Voice-Pro dependencies without the GPU libraries?

Yes. The repository uses optional dependency groups. Installing without the `gpu` extra will pull **torch**, **torchvision**, and **torchaudio** CPU wheels plus **onnxruntime** instead of **onnxruntime-gpu**, significantly reducing installation size while maintaining full functionality on CPU-only systems.