Voice-Pro Application Directory Structure: A Complete Guide to the Modular Python Architecture

Voice-Pro organizes its codebase into a three-layer Python architecture separating the Gradio UI, controller logic, and core processing engines across distinct modules in the app/ directory.

Voice-Pro is an open-source AI voice processing application built by ABUS AI Korea. Understanding the Voice-Pro application directory structure is essential for developers who want to extend functionality, debug processing pipelines, or deploy custom voice synthesis workflows.

Top-Level Directory Layout

The repository root maintains a clean separation between application code, configuration, and runtime assets. According to the abus-aikorea/voice-pro source code, the top-level entries follow this organization:

  • app/ – All application source code, divided into UI tabs, controllers, and core engines
  • installer_files/ – Runtime-generated directory containing the UV-managed Python environment and portable FFmpeg binaries (not version-controlled)
  • start-voice.py – Primary entry point for the Voice product variant
  • start-abus.py – General launcher supporting multiple product variants
  • one_click.py – Installation sanity check and repair script
  • pyproject.toml & uv.lock – Dependency declarations and locked versions for reproducible installs
  • .env.example – Template for optional Azure credentials (translation/TTS services)
  • README.md – High-level documentation and quick-start instructions

The Three-Layer Architecture Inside app/

The app/ directory implements a strict separation of concerns through three distinct layers, making it easy to swap ASR engines, TTS models, or translation backends without modifying UI code.

UI Layer: tab_*.py Modules

These files define the Gradio interface components, widget layouts, and event wiring for each feature tab:

  • tab_gulliver.py – Implements the "Dubbing Studio" tab that chains download → speech-to-text → translate → TTS pipelines
  • tab_translate.py – UI for batch translation of video files
  • tab_subtitle.py – Interface for generating Whisper subtitles

Each tab module instantiates Gradio components and binds them to controller methods.

Controller Layer: gradio_*.py Classes

Controllers hold UI state and orchestrate the processing pipeline, acting as intermediaries between the interface and core engines:

These classes expose methods called by tab event handlers and manage the flow of data between user inputs and processing engines.

Core Processing Layer: abus_*.py Engines

The core modules contain Gradio-free business logic for heavy AI/ML operations:

This layer handles model downloading, inference execution, and file I/O operations.

Configuration and Runtime Assets

Voice-Pro persists user settings and model manifests within the app/ directory:

  • config-user.json5 – User-editable configuration storing UI settings, engine selections, and default parameters
  • abus_hf_files-*.json – Manifest files listing Hugging Face model assets to download on first run (e.g., abus_hf_files-voice.json)
  • workspace/ – Runtime output directory for processed audio, translations, and temporary files (managed via abus_path.py)
  • model/ – Downloaded model weights cached locally for offline inference

Entry Points and Launchers

The application initialization flow begins at the root level scripts:

start-voice.py initializes the UV environment, loads config-user.json5, and calls abus_app_voice.create_ui() to build the Gradio interface:


# start-voice.py

import sys
from app.abus_app_voice import create_ui

if __name__ == "__main__":
    create_ui()  # Builds Gradio UI and starts the server

start-abus.py serves as a general launcher used by batch scripts to select between product variants (voice, upscaler, etc.).

Dependency Management

The project uses uv for fast, reproducible Python environment management:

  • pyproject.toml – Declares required packages including Gradio, PyTorch, and audio processing libraries, with optional extras for gpu or cpu installations
  • uv.lock – Pins exact dependency versions ensuring consistent builds across environments

The installer_files/ directory is generated at runtime by the configure.* scripts, housing the self-contained Python interpreter and portable FFmpeg binaries when system FFmpeg is unavailable.

Practical Code Examples

Running ASR Directly

You can bypass the UI and call core modules directly for batch processing:

from app.abus_asr_faster_whisper import FasterWhisperASR

asr = FasterWhisperASR(model_name="large-v2")
transcript = asr.transcribe("workspace/input/audio.wav")
print(transcript)

Switching Translation Backends

Runtime backend selection allows flexible deployment with or without cloud credentials:

from app.abus_translate_azure import AzureTranslator
from app.abus_translate_deep import DeepTranslator

# Choose based on user config (DeepTranslator is the default)

translator = AzureTranslator() if use_azure else DeepTranslator()
result = translator.translate("안녕하세요", target_lang="en")
print(result)  # → "Hello"

Extending with a New Tab

Adding functionality follows the established three-layer pattern:


# In app/tab_myfeature.py

import gradio as gr
from app.gradio_myfeature import GradioMyFeature

def my_feature_tab():
    with gr.Blocks() as demo:
        btn = gr.Button("Run my feature")
        out = gr.Textbox()
        btn.click(fn=GradioMyFeature().run, inputs=None, outputs=out)
    return demo

Summary

  • Voice-Pro organizes code into three layers: UI (tab_*.py), Controller (gradio_*.py), and Core (abus_*.py), enabling modular AI engine swaps
  • The app/ directory contains all business logic, while installer_files/ (runtime-generated) houses the Python environment
  • Entry points (start-voice.py, start-abus.py) initialize the UV-managed environment before loading the Gradio interface via abus_app_voice.create_ui()
  • Configuration persists in config-user.json5, with model manifests in abus_hf_files-*.json and outputs directed to workspace/
  • Dependencies are locked via uv.lock and declared in pyproject.toml, supporting both CPU and GPU extras

Frequently Asked Questions

Where are the AI model weights stored in Voice-Pro?

Downloaded model weights are cached in the model/ directory under the application root, with download manifests defined in app/abus_hf_files-*.json files. The abus_path.py module provides centralized path resolution for both model storage and the workspace/ output directory.

How does Voice-Pro separate the UI from business logic?

The architecture enforces strict separation through the tab_*.py UI layer (Gradio widgets), the gradio_*.py controller layer (state management and orchestration), and the abus_*.py core layer (Gradio-free processing). Controllers like GradioGulliver in app/gradio_gulliver.py wire UI events to core engines such as abus_asr_faster_whisper.py.

What is the purpose of the installer_files directory?

The installer_files/ directory is generated at runtime by the configuration scripts and contains the UV-managed Python environment along with portable FFmpeg binaries when the system PATH lacks FFmpeg. This directory is not version-controlled and ensures the application runs with isolated, reproducible dependencies.

Which file should I modify to change the default ASR engine?

To modify ASR behavior, edit app/abus_asr_faster_whisper.py (the core engine) or app/gradio_gulliver.py (the controller that instantiates the ASR class). The tab-specific UI logic resides in app/tab_gulliver.py, but engine-specific parameters and model selection are managed in the controller and core layers.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →