# Voice-Pro Voice Cloning Technologies: A Complete Guide to 5 Integrated TTS Engines

> Discover Voice-Pro's five integrated voice cloning technologies including CosyVoice, RVC, F5-TTS/E2-TTS, Microsoft Edge TTS, and Azure Speech Service. Explore this comprehensive guide.

- Repository: [ABUS/voice-pro](https://github.com/abus-aikorea/voice-pro)
- Tags: deep-dive
- Published: 2026-08-03

---

**Voice-Pro integrates five distinct voice cloning technologies—CosyVoice (CosyVoice2-0.5B and Fun-CosyVoice-3-0.5B), RVC for real-time voice conversion, F5-TTS/E2-TTS diffusion models, Microsoft Edge TTS, and Azure Speech Service—each implemented as dedicated Python classes within the `app/` directory.**

Voice-Pro by abus-aikorea is an open-source application that bundles multiple state-of-the-art voice cloning and text-to-speech engines into a unified Gradio interface. According to the source code repository, these technologies are wrapped in consistent Python controllers that handle everything from reference audio preprocessing to post-processing effects, enabling seamless switching between local and cloud-based voice synthesis.

## Core Voice Cloning Technologies in Voice-Pro

The application implements five primary TTS and voice conversion engines, each residing in specific modules under the `app/` folder.

### CosyVoice with Zero-Shot and Cross-Lingual Support

The **CosyVoice** integration provides high-quality zero-shot, cross-lingual, and instruction-guided TTS through the `CosyVoiceInference` class in [`app/abus_tts_cosyvoice.py`](https://github.com/abus-aikorea/voice-pro/blob/main/app/abus_tts_cosyvoice.py). This implementation supports two model variants: **CosyVoice2-0.5B** for generic multilingual synthesis and **Fun-CosyVoice-3-0.5B** optimized for Korean-centric applications.

Key capabilities include three distinct inference modes—`"Zero-Shot"`, `"Cross-Lingual"`, and `"Instruct"`—accessible via the `request_tts()` method. The class handles automatic model downloading, prompt preparation, and speed adjustment through the `speed_factor` parameter.

### RVC (Real-Time Voice Conversion)

**RVC** (Retrieval-based Voice Conversion) alters speaker identity while preserving linguistic content. The core inference pipeline lives in [`app/abus_rvc.py`](https://github.com/abus-aikorea/voice-pro/blob/main/app/abus_rvc.py) within the `RVC` class, which manages model folders, prerequisite downloads, and the low-level `infer_pipeline` execution.

For end-to-end TTS workflows, Voice-Pro provides the `TTSRVC` wrapper class in [`app/abus_tts_rvc.py`](https://github.com/abus-aikorea/voice-pro/blob/main/app/abus_tts_rvc.py). This wrapper combines standard TTS generation with RVC post-processing, allowing any Azure or Edge TTS output to be converted into a target voice style using the `infer()` method.

### F5-TTS and E2-TTS Diffusion Models

The **F5-TTS** integration implements diffusion-based text-to-speech synthesis via the `F5TTS` class in [`app/abus_tts_f5.py`](https://github.com/abus-aikorea/voice-pro/blob/main/app/abus_tts_f5.py). This controller supports both the **SWivid/F5-TTS_v1** and **SWivid/E2-TTS** model families, selectable through the `select_model()` method.

Users configure synthesis parameters including `speed_factor` and `model_choice`, then generate audio via `infer_single()`, which processes the reference celebrity audio and transcript through the diffusion inference pipeline.

### Microsoft Edge TTS (Free Local Option)

**EdgeTTS** in [`app/abus_tts_edge.py`](https://github.com/abus-aikorea/voice-pro/blob/main/app/abus_tts_edge.py) provides free, locally-executed text-to-speech using the Microsoft Edge browser's speech synthesis API. This class serves as the default fallback when Azure credentials are absent, requiring no cloud subscription or API keys.

The implementation exposes a `text_to_voice()` method that accepts standard parameters including `voice_name` (e.g., `"en-US-GuyNeural"`), `speed_factor`, and `semitones` for pitch adjustment.

### Azure Speech Service TTS (Cloud-Based)

The **AzureTTS** class in [`app/abus_tts_azure.py`](https://github.com/abus-aikorea/voice-pro/blob/main/app/abus_tts_azure.py) integrates Microsoft Azure's Speech Service for high-fidelity cloud-based synthesis. Voice-Pro automatically activates this engine when a valid Azure subscription key is detected, offering superior voice quality compared to the local Edge alternative.

## The Voice Cloning Pipeline Architecture

All five engines share a standardized three-stage workflow implemented across the codebase:

1. **Reference Audio Extraction**: The system processes celebrity reference clips and transcripts via `preprocess_ref_audio_text` to create prompt wav files compatible with CosyVoice and F5-TTS.

2. **TTS Generation**: The selected controller—whether `CosyVoiceInference`, `F5TTS`, `AzureTTS`, or `EdgeTTS`—synthesizes speech for each input line or subtitle segment.

3. **Post-Processing**: Output undergoes silence trimming via `AbusAudio.trim_silence_file`, optional voice conversion through the RVC pipeline, and final stereo conversion using `ffmpeg_to_stereo`.

## Implementation Examples for Each Voice Cloning Engine

The following snippets demonstrate direct programmatic access to each engine, mirroring the internal calls made by Voice-Pro's Gradio UI controllers.

### CosyVoice Inference

```python
from app.abus_tts_cosyvoice import CosyVoiceInference

cosy = CosyVoiceInference(model_name="CosyVoice2-0.5B")
cosy.request_tts(
    line="Hello, world!",
    output_file="/tmp/output.wav",
    ref_audio="/path/to/celebrity.wav",
    ref_text="Sample transcript for the reference voice.",
    inference_mode="Zero-Shot",
    speed_factor=1.0,
    audio_format="wav"
)

```

### RVC Voice Conversion

```python
from app.abus_tts_rvc import TTSRVC

tts = TTSRVC()
tts_audio, rvc_audio = tts.infer(
    text="This is a test sentence.",
    tts_voice="en-US-AriaNeural",
    semitones=0,
    speed_factor=1.0,
    volume_factor=1.0,
    audio_format="wav",
    rvc_voice="my_rvc_voice"
)

```

### F5-TTS Diffusion Synthesis

```python
from app.abus_tts_f5 import F5TTS

f5 = F5TTS()
f5.select_model("SWivid/F5-TTS_v1")
f5.infer_single(
    dubbing_text="A quick brown fox jumps over the lazy dog.",
    output_file="/tmp/f5_output.wav",
    celeb_audio="/path/to/reference.wav",
    celeb_transcript="Reference transcript.",
    model_choice="SWivid/F5-TTS_v1",
    speed_factor=1.0,
    audio_format="wav"
)

```

### Edge and Azure TTS

```python
from app.abus_tts_edge import EdgeTTS
from app.abus_tts_azure import AzureTTS
from app.abus_genuine import azure_text_api_working

tts_engine = AzureTTS() if azure_text_api_working() else EdgeTTS()
tts_engine.text_to_voice(
    text="Testing the Edge TTS engine.",
    output_file="/tmp/edge_output.wav",
    voice_name="en-US-GuyNeural",
    semitones=0,
    speed_factor=1.0,
    volume_factor=1.0,
    audio_format="wav"
)

```

## Summary

- **Voice-Pro integrates five voice cloning technologies**: CosyVoice (CosyVoice2-0.5B and Fun-CosyVoice-3-0.5B), RVC for real-time conversion, F5-TTS/E2-TTS diffusion models, Microsoft Edge TTS, and Azure Speech Service.
- **Each engine is encapsulated in a dedicated class**: `CosyVoiceInference`, `RVC`/`TTSRVC`, `F5TTS`, `EdgeTTS`, and `AzureTTS`, all located in the `app/` directory with standardized `request_tts` or `infer_single` interfaces.
- **RVC functions as a post-processor**: Converting any TTS output into a target voice identity while preserving the original content.
- **Edge TTS serves as the free default**: Automatically falling back when Azure credentials are unavailable.
- **The unified pipeline** handles reference audio extraction, synthesis, and post-processing (silence trimming, stereo conversion) consistently across all engines.

## Frequently Asked Questions

### What is the difference between CosyVoice and F5-TTS in Voice-Pro?

**CosyVoice** provides zero-shot voice cloning with explicit support for cross-lingual synthesis and instruction following through the `CosyVoiceInference` class, while **F5-TTS** utilizes diffusion-based generation via the `F5TTS` class in [`app/abus_tts_f5.py`](https://github.com/abus-aikorea/voice-pro/blob/main/app/abus_tts_f5.py), offering different model architectures (SWivid/F5-TTS_v1 vs SWivid/E2-TTS) with configurable speed parameters. CosyVoice excels at maintaining speaker similarity across languages, whereas F5-TTS focuses on high-quality diffusion-based audio generation.

### How does RVC integrate with other TTS engines?

The `TTSRVC` wrapper class in [`app/abus_tts_rvc.py`](https://github.com/abus-aikorea/voice-pro/blob/main/app/abus_tts_rvc.py) chains standard TTS generation with RVC voice conversion, accepting outputs from Edge or Azure TTS and processing them through the `RVC` class's `infer_pipeline` in [`app/abus_rvc.py`](https://github.com/abus-aikorea/voice-pro/blob/main/app/abus_rvc.py) to alter speaker identity while preserving linguistic content. This allows users to convert generic TTS voices into custom trained voices without modifying the underlying synthesis engine.

### Can I use Voice-Pro without an Azure subscription?

Yes. When the `azure_text_api_working()` check fails or no valid Azure key is detected, Voice-Pro automatically defaults to the `EdgeTTS` class in [`app/abus_tts_edge.py`](https://github.com/abus-aikorea/voice-pro/blob/main/app/abus_tts_edge.py), which uses the free Microsoft Edge browser speech synthesis API locally without requiring cloud credentials or API keys.

### Which voice cloning technology supports Korean language best?

**Fun-CosyVoice-3-0.5B**, accessible through the `CosyVoiceInference` class by selecting the appropriate model name, is specifically optimized for Korean-centric applications according to the source code implementation in [`app/abus_tts_cosyvoice.py`](https://github.com/abus-aikorea/voice-pro/blob/main/app/abus_tts_cosyvoice.py), while also supporting cross-lingual capabilities for other languages.