How Voice-Pro Selects and Uses ASR Engines: Factory Pattern and Configuration Guide
Voice-Pro uses a configuration-driven factory pattern to switch between Faster-Whisper, Whisper, and Whisper-Timestamped engines, storing the user’s choice in app/config-user.json5 and defaulting to Faster-Whisper for optimal performance.
Voice-Pro supports multiple Automatic Speech Recognition (ASR) engines to balance accuracy, speed, and timestamp precision. The selection mechanism is implemented in the Gradio controller classes (GradioGulliver and GradioASR) and uses a factory method to instantiate the appropriate inference wrapper. This article examines how the abus-aikorea/voice-pro repository handles ASR engine configuration, instantiation, and runtime execution.
User Configuration and Default Engine Selection
The ASR engine selection persists across sessions through a user-specific configuration file.
Reading the Configuration
When the UI initializes, the controller reads the asr_engine key from app/config-user.json5. If the key is absent, the system defaults to 'faster-whisper':
asr_engine = self.user_config.get("asr_engine", 'faster-whisper')
self.whisper_inf = self.switch_case(asr_engine)
This logic appears in both app/gradio_gulliver.py (lines 59-61) and app/gradio_asr.py (lines 29-31), ensuring consistent behavior across the full dubbing pipeline and the dedicated ASR tab.
Persisting User Preferences
After a user selects an engine and model, the controller writes the values back to the configuration:
self.user_config.set("asr_engine", asr_engine)
self.user_config.set(f'{asr_engine.replace("-", "_")}_model', modelName)
This persistence mechanism, found in app/gradio_gulliver.py (lines 71-74), ensures subsequent sessions initialize with the previously selected ASR engine.
The Factory Pattern Implementation
Voice-Pro decouples engine selection from instantiation using a private factory method that maps string identifiers to concrete inference classes.
The switch_case Method
Both GradioGulliver and GradioASR implement a switch_case method that acts as a factory:
def switch_case(self, case):
switch_dict = {
'faster-whisper': lambda: FasterWhisperInference(),
'whisper': lambda: WhisperInference(),
'whisper-timestamped': lambda: WhisperTimestampedInference()
}
return switch_dict.get(case, lambda: FasterWhisperInference())()
This implementation appears in app/gradio_gulliver.py (lines 67-73) and app/gradio_asr.py (lines 36-42). The method returns a lambda instantiation of the appropriate wrapper class, defaulting to FasterWhisperInference if an invalid identifier is provided.
ASR Engine Implementations
Voice-Pro provides three specialized wrapper classes that expose a unified interface for model discovery and transcription.
Faster-Whisper Wrapper
The app/abus_asr_faster_whisper.py module wraps the Faster-Whisper model, offering optimized inference with methods such as available_models(), available_langs(), and transcribe_file().
Original Whisper Wrapper
Located in app/abus_asr_whisper.py, this wrapper provides a thin interface to the original OpenAI Whisper model, maintaining the same method signatures for drop-in compatibility.
Whisper-Timestamped Wrapper
The app/abus_asr_whisper_timestamped.py module extends the base Whisper functionality to provide word-level timestamp alignment, essential for precise subtitle generation.
Unified API Interface
All three implementations expose a consistent interface:
available_models()– Returns a list of supported model sizesavailable_langs()– Returns supported language codestranscribe_file()– Executes speech-to-text conversion
This uniform API allows the Gradio controllers to treat any engine interchangeably without conditional logic scattered throughout the codebase.
Runtime Transcription Workflow
When a user initiates transcription, the controller instantiates the selected engine and executes the inference pipeline.
Engine Instantiation
The controller re-instantiates the inference class whenever the engine selection changes:
self.whisper_inf = self.switch_case(asr_engine)
This ensures the correct backend is loaded with the appropriate model weights and compute settings.
Executing Transcription
The transcription workflow in app/gradio_gulliver.py (lines 85-91) demonstrates the unified API in action:
self.whisper_inf = self.switch_case(asr_engine)
subtitles = self.whisper_inf.transcribe_file(input_path, params, False, gr.Progress())
The transcribe_file method accepts the audio path, parameters (model size, language, compute type), a boolean flag, and a Gradio progress tracker.
Complete Usage Example
To run transcription programmatically with a specific engine:
params = WhisperParameters(
model_size="large",
lang="english",
compute_type="float16"
)
audio_path = "/workspace/audio.wav"
subtitles = self.whisper_inf.transcribe_file(audio_path, params, False, gr.Progress())
Dynamic UI Updates
The interface adapts dynamically when users switch engines, updating available models and languages accordingly.
Refreshing Model Lists
When the engine selector changes, the controller calls update_whisper_models() to repopulate the model dropdown:
def update_whisper_models(self, asr_engine):
whisper_inf = self.switch_case(asr_engine)
model_list = whisper_inf.available_models()
# Returns Gradio update object with new choices
This method, found in app/gradio_gulliver.py (lines 86-93), creates a temporary instance of the selected engine to query its available_models() method, ensuring the UI only presents compatible options.
Summary
- Voice-Pro supports three ASR engines: Faster-Whisper, Whisper, and Whisper-Timestamped, configurable via
app/config-user.json5. - The
switch_casefactory method inGradioGulliverandGradioASRmaps string identifiers to concrete inference classes, defaulting to Faster-Whisper. - Each engine wrapper (
app/abus_asr_faster_whisper.py,app/abus_asr_whisper.py,app/abus_asr_whisper_timestamped.py) implements a unified API for model discovery and transcription. - Runtime transcription uses the
transcribe_file()method with consistent parameters regardless of the selected backend. - User selections persist across sessions through the
UserConfigclass, and the UI dynamically updates model lists based on the active engine.
Frequently Asked Questions
How do I switch ASR engines in Voice-Pro?
Navigate to the engine dropdown in the Gradio interface and select your preferred backend. The controller stores this choice in app/config-user.json5 under the asr_engine key, then reinstantiates the inference class via the switch_case factory method.
What is the default ASR engine in Voice-Pro?
The system defaults to Faster-Whisper if no previous selection exists. This default is hardcoded in the user_config.get("asr_engine", 'faster-whisper') call found in both app/gradio_gulliver.py and app/gradio_asr.py.
Can I use Whisper-Timestamped for word-level alignment?
Yes. Select Whisper-Timestamped from the engine dropdown. The app/abus_asr_whisper_timestamped.py wrapper extends the base Whisper functionality to provide precise timestamped output suitable for subtitle generation.
Where is the ASR engine selection logic implemented?
The core selection logic resides in app/gradio_gulliver.py (for the full dubbing pipeline) and app/gradio_asr.py (for the ASR-only tab). Both files contain the switch_case factory method and configuration persistence code.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →