Voice-Pro Application Directory Structure: A Complete Guide to the Modular Python Architecture
Voice-Pro organizes its codebase into a three-layer Python architecture separating the Gradio UI, controller logic, and core processing engines across distinct modules in the app/ directory.
Voice-Pro is an open-source AI voice processing application built by ABUS AI Korea. Understanding the Voice-Pro application directory structure is essential for developers who want to extend functionality, debug processing pipelines, or deploy custom voice synthesis workflows.
Top-Level Directory Layout
The repository root maintains a clean separation between application code, configuration, and runtime assets. According to the abus-aikorea/voice-pro source code, the top-level entries follow this organization:
app/– All application source code, divided into UI tabs, controllers, and core enginesinstaller_files/– Runtime-generated directory containing the UV-managed Python environment and portable FFmpeg binaries (not version-controlled)start-voice.py– Primary entry point for the Voice product variantstart-abus.py– General launcher supporting multiple product variantsone_click.py– Installation sanity check and repair scriptpyproject.toml&uv.lock– Dependency declarations and locked versions for reproducible installs.env.example– Template for optional Azure credentials (translation/TTS services)README.md– High-level documentation and quick-start instructions
The Three-Layer Architecture Inside app/
The app/ directory implements a strict separation of concerns through three distinct layers, making it easy to swap ASR engines, TTS models, or translation backends without modifying UI code.
UI Layer: tab_*.py Modules
These files define the Gradio interface components, widget layouts, and event wiring for each feature tab:
tab_gulliver.py– Implements the "Dubbing Studio" tab that chains download → speech-to-text → translate → TTS pipelinestab_translate.py– UI for batch translation of video filestab_subtitle.py– Interface for generating Whisper subtitles
Each tab module instantiates Gradio components and binds them to controller methods.
Controller Layer: gradio_*.py Classes
Controllers hold UI state and orchestrate the processing pipeline, acting as intermediaries between the interface and core engines:
gradio_gulliver.py– Contains theGradioGulliverclass that coordinates the full dubbing workflowgradio_translate.py– Manages batch translation workflowsgradio_tts_cosyvoice.py– Handles CosyVoice TTS-specific UI state and inference calls
These classes expose methods called by tab event handlers and manage the flow of data between user inputs and processing engines.
Core Processing Layer: abus_*.py Engines
The core modules contain Gradio-free business logic for heavy AI/ML operations:
abus_asr_faster_whisper.py– Faster-Whisper ASR engine implementationabus_tts_cosyvoice.py– CosyVoice TTS backend with voice cloning supportabus_translate_deep.py– Free DeepL-based translator (fallback when Azure credentials are absent)abus_translate_azure.py– Azure Cognitive Services translation integrationabus_path.py– Centralized path helpers forworkspace/,model/, and temporary directories
This layer handles model downloading, inference execution, and file I/O operations.
Configuration and Runtime Assets
Voice-Pro persists user settings and model manifests within the app/ directory:
config-user.json5– User-editable configuration storing UI settings, engine selections, and default parametersabus_hf_files-*.json– Manifest files listing Hugging Face model assets to download on first run (e.g.,abus_hf_files-voice.json)workspace/– Runtime output directory for processed audio, translations, and temporary files (managed viaabus_path.py)model/– Downloaded model weights cached locally for offline inference
Entry Points and Launchers
The application initialization flow begins at the root level scripts:
start-voice.py initializes the UV environment, loads config-user.json5, and calls abus_app_voice.create_ui() to build the Gradio interface:
# start-voice.py
import sys
from app.abus_app_voice import create_ui
if __name__ == "__main__":
create_ui() # Builds Gradio UI and starts the server
start-abus.py serves as a general launcher used by batch scripts to select between product variants (voice, upscaler, etc.).
Dependency Management
The project uses uv for fast, reproducible Python environment management:
pyproject.toml– Declares required packages including Gradio, PyTorch, and audio processing libraries, with optional extras forgpuorcpuinstallationsuv.lock– Pins exact dependency versions ensuring consistent builds across environments
The installer_files/ directory is generated at runtime by the configure.* scripts, housing the self-contained Python interpreter and portable FFmpeg binaries when system FFmpeg is unavailable.
Practical Code Examples
Running ASR Directly
You can bypass the UI and call core modules directly for batch processing:
from app.abus_asr_faster_whisper import FasterWhisperASR
asr = FasterWhisperASR(model_name="large-v2")
transcript = asr.transcribe("workspace/input/audio.wav")
print(transcript)
Switching Translation Backends
Runtime backend selection allows flexible deployment with or without cloud credentials:
from app.abus_translate_azure import AzureTranslator
from app.abus_translate_deep import DeepTranslator
# Choose based on user config (DeepTranslator is the default)
translator = AzureTranslator() if use_azure else DeepTranslator()
result = translator.translate("안녕하세요", target_lang="en")
print(result) # → "Hello"
Extending with a New Tab
Adding functionality follows the established three-layer pattern:
# In app/tab_myfeature.py
import gradio as gr
from app.gradio_myfeature import GradioMyFeature
def my_feature_tab():
with gr.Blocks() as demo:
btn = gr.Button("Run my feature")
out = gr.Textbox()
btn.click(fn=GradioMyFeature().run, inputs=None, outputs=out)
return demo
Summary
- Voice-Pro organizes code into three layers: UI (
tab_*.py), Controller (gradio_*.py), and Core (abus_*.py), enabling modular AI engine swaps - The
app/directory contains all business logic, whileinstaller_files/(runtime-generated) houses the Python environment - Entry points (
start-voice.py,start-abus.py) initialize the UV-managed environment before loading the Gradio interface viaabus_app_voice.create_ui() - Configuration persists in
config-user.json5, with model manifests inabus_hf_files-*.jsonand outputs directed toworkspace/ - Dependencies are locked via
uv.lockand declared inpyproject.toml, supporting both CPU and GPU extras
Frequently Asked Questions
Where are the AI model weights stored in Voice-Pro?
Downloaded model weights are cached in the model/ directory under the application root, with download manifests defined in app/abus_hf_files-*.json files. The abus_path.py module provides centralized path resolution for both model storage and the workspace/ output directory.
How does Voice-Pro separate the UI from business logic?
The architecture enforces strict separation through the tab_*.py UI layer (Gradio widgets), the gradio_*.py controller layer (state management and orchestration), and the abus_*.py core layer (Gradio-free processing). Controllers like GradioGulliver in app/gradio_gulliver.py wire UI events to core engines such as abus_asr_faster_whisper.py.
What is the purpose of the installer_files directory?
The installer_files/ directory is generated at runtime by the configuration scripts and contains the UV-managed Python environment along with portable FFmpeg binaries when the system PATH lacks FFmpeg. This directory is not version-controlled and ensures the application runs with isolated, reproducible dependencies.
Which file should I modify to change the default ASR engine?
To modify ASR behavior, edit app/abus_asr_faster_whisper.py (the core engine) or app/gradio_gulliver.py (the controller that instantiates the ASR class). The tab-specific UI logic resides in app/tab_gulliver.py, but engine-specific parameters and model selection are managed in the controller and core layers.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →