Voice-Pro Voice Cloning Technologies: A Complete Guide to 5 Integrated TTS Engines
Voice-Pro integrates five distinct voice cloning technologies—CosyVoice (CosyVoice2-0.5B and Fun-CosyVoice-3-0.5B), RVC for real-time voice conversion, F5-TTS/E2-TTS diffusion models, Microsoft Edge TTS, and Azure Speech Service—each implemented as dedicated Python classes within the app/ directory.
Voice-Pro by abus-aikorea is an open-source application that bundles multiple state-of-the-art voice cloning and text-to-speech engines into a unified Gradio interface. According to the source code repository, these technologies are wrapped in consistent Python controllers that handle everything from reference audio preprocessing to post-processing effects, enabling seamless switching between local and cloud-based voice synthesis.
Core Voice Cloning Technologies in Voice-Pro
The application implements five primary TTS and voice conversion engines, each residing in specific modules under the app/ folder.
CosyVoice with Zero-Shot and Cross-Lingual Support
The CosyVoice integration provides high-quality zero-shot, cross-lingual, and instruction-guided TTS through the CosyVoiceInference class in app/abus_tts_cosyvoice.py. This implementation supports two model variants: CosyVoice2-0.5B for generic multilingual synthesis and Fun-CosyVoice-3-0.5B optimized for Korean-centric applications.
Key capabilities include three distinct inference modes—"Zero-Shot", "Cross-Lingual", and "Instruct"—accessible via the request_tts() method. The class handles automatic model downloading, prompt preparation, and speed adjustment through the speed_factor parameter.
RVC (Real-Time Voice Conversion)
RVC (Retrieval-based Voice Conversion) alters speaker identity while preserving linguistic content. The core inference pipeline lives in app/abus_rvc.py within the RVC class, which manages model folders, prerequisite downloads, and the low-level infer_pipeline execution.
For end-to-end TTS workflows, Voice-Pro provides the TTSRVC wrapper class in app/abus_tts_rvc.py. This wrapper combines standard TTS generation with RVC post-processing, allowing any Azure or Edge TTS output to be converted into a target voice style using the infer() method.
F5-TTS and E2-TTS Diffusion Models
The F5-TTS integration implements diffusion-based text-to-speech synthesis via the F5TTS class in app/abus_tts_f5.py. This controller supports both the SWivid/F5-TTS_v1 and SWivid/E2-TTS model families, selectable through the select_model() method.
Users configure synthesis parameters including speed_factor and model_choice, then generate audio via infer_single(), which processes the reference celebrity audio and transcript through the diffusion inference pipeline.
Microsoft Edge TTS (Free Local Option)
EdgeTTS in app/abus_tts_edge.py provides free, locally-executed text-to-speech using the Microsoft Edge browser's speech synthesis API. This class serves as the default fallback when Azure credentials are absent, requiring no cloud subscription or API keys.
The implementation exposes a text_to_voice() method that accepts standard parameters including voice_name (e.g., "en-US-GuyNeural"), speed_factor, and semitones for pitch adjustment.
Azure Speech Service TTS (Cloud-Based)
The AzureTTS class in app/abus_tts_azure.py integrates Microsoft Azure's Speech Service for high-fidelity cloud-based synthesis. Voice-Pro automatically activates this engine when a valid Azure subscription key is detected, offering superior voice quality compared to the local Edge alternative.
The Voice Cloning Pipeline Architecture
All five engines share a standardized three-stage workflow implemented across the codebase:
-
Reference Audio Extraction: The system processes celebrity reference clips and transcripts via
preprocess_ref_audio_textto create prompt wav files compatible with CosyVoice and F5-TTS. -
TTS Generation: The selected controller—whether
CosyVoiceInference,F5TTS,AzureTTS, orEdgeTTS—synthesizes speech for each input line or subtitle segment. -
Post-Processing: Output undergoes silence trimming via
AbusAudio.trim_silence_file, optional voice conversion through the RVC pipeline, and final stereo conversion usingffmpeg_to_stereo.
Implementation Examples for Each Voice Cloning Engine
The following snippets demonstrate direct programmatic access to each engine, mirroring the internal calls made by Voice-Pro's Gradio UI controllers.
CosyVoice Inference
from app.abus_tts_cosyvoice import CosyVoiceInference
cosy = CosyVoiceInference(model_name="CosyVoice2-0.5B")
cosy.request_tts(
line="Hello, world!",
output_file="/tmp/output.wav",
ref_audio="/path/to/celebrity.wav",
ref_text="Sample transcript for the reference voice.",
inference_mode="Zero-Shot",
speed_factor=1.0,
audio_format="wav"
)
RVC Voice Conversion
from app.abus_tts_rvc import TTSRVC
tts = TTSRVC()
tts_audio, rvc_audio = tts.infer(
text="This is a test sentence.",
tts_voice="en-US-AriaNeural",
semitones=0,
speed_factor=1.0,
volume_factor=1.0,
audio_format="wav",
rvc_voice="my_rvc_voice"
)
F5-TTS Diffusion Synthesis
from app.abus_tts_f5 import F5TTS
f5 = F5TTS()
f5.select_model("SWivid/F5-TTS_v1")
f5.infer_single(
dubbing_text="A quick brown fox jumps over the lazy dog.",
output_file="/tmp/f5_output.wav",
celeb_audio="/path/to/reference.wav",
celeb_transcript="Reference transcript.",
model_choice="SWivid/F5-TTS_v1",
speed_factor=1.0,
audio_format="wav"
)
Edge and Azure TTS
from app.abus_tts_edge import EdgeTTS
from app.abus_tts_azure import AzureTTS
from app.abus_genuine import azure_text_api_working
tts_engine = AzureTTS() if azure_text_api_working() else EdgeTTS()
tts_engine.text_to_voice(
text="Testing the Edge TTS engine.",
output_file="/tmp/edge_output.wav",
voice_name="en-US-GuyNeural",
semitones=0,
speed_factor=1.0,
volume_factor=1.0,
audio_format="wav"
)
Summary
- Voice-Pro integrates five voice cloning technologies: CosyVoice (CosyVoice2-0.5B and Fun-CosyVoice-3-0.5B), RVC for real-time conversion, F5-TTS/E2-TTS diffusion models, Microsoft Edge TTS, and Azure Speech Service.
- Each engine is encapsulated in a dedicated class:
CosyVoiceInference,RVC/TTSRVC,F5TTS,EdgeTTS, andAzureTTS, all located in theapp/directory with standardizedrequest_ttsorinfer_singleinterfaces. - RVC functions as a post-processor: Converting any TTS output into a target voice identity while preserving the original content.
- Edge TTS serves as the free default: Automatically falling back when Azure credentials are unavailable.
- The unified pipeline handles reference audio extraction, synthesis, and post-processing (silence trimming, stereo conversion) consistently across all engines.
Frequently Asked Questions
What is the difference between CosyVoice and F5-TTS in Voice-Pro?
CosyVoice provides zero-shot voice cloning with explicit support for cross-lingual synthesis and instruction following through the CosyVoiceInference class, while F5-TTS utilizes diffusion-based generation via the F5TTS class in app/abus_tts_f5.py, offering different model architectures (SWivid/F5-TTS_v1 vs SWivid/E2-TTS) with configurable speed parameters. CosyVoice excels at maintaining speaker similarity across languages, whereas F5-TTS focuses on high-quality diffusion-based audio generation.
How does RVC integrate with other TTS engines?
The TTSRVC wrapper class in app/abus_tts_rvc.py chains standard TTS generation with RVC voice conversion, accepting outputs from Edge or Azure TTS and processing them through the RVC class's infer_pipeline in app/abus_rvc.py to alter speaker identity while preserving linguistic content. This allows users to convert generic TTS voices into custom trained voices without modifying the underlying synthesis engine.
Can I use Voice-Pro without an Azure subscription?
Yes. When the azure_text_api_working() check fails or no valid Azure key is detected, Voice-Pro automatically defaults to the EdgeTTS class in app/abus_tts_edge.py, which uses the free Microsoft Edge browser speech synthesis API locally without requiring cloud credentials or API keys.
Which voice cloning technology supports Korean language best?
Fun-CosyVoice-3-0.5B, accessible through the CosyVoiceInference class by selecting the appropriate model name, is specifically optimized for Korean-centric applications according to the source code implementation in app/abus_tts_cosyvoice.py, while also supporting cross-lingual capabilities for other languages.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →