# How Voice-Pro Selects and Uses ASR Engines: Factory Pattern and Configuration Guide

> Learn how Voice-Pro selects and uses ASR engines with its factory pattern and configuration. Discover efficient switching between Whisper variants for optimal performance.

- Repository: [ABUS/voice-pro](https://github.com/abus-aikorea/voice-pro)
- Tags: how-to-guide
- Published: 2026-08-03

---

**Voice-Pro uses a configuration-driven factory pattern to switch between Faster-Whisper, Whisper, and Whisper-Timestamped engines, storing the user’s choice in `app/config-user.json5` and defaulting to Faster-Whisper for optimal performance.**

Voice-Pro supports multiple Automatic Speech Recognition (ASR) engines to balance accuracy, speed, and timestamp precision. The selection mechanism is implemented in the Gradio controller classes (`GradioGulliver` and `GradioASR`) and uses a factory method to instantiate the appropriate inference wrapper. This article examines how the abus-aikorea/voice-pro repository handles ASR engine configuration, instantiation, and runtime execution.

## User Configuration and Default Engine Selection

The ASR engine selection persists across sessions through a user-specific configuration file.

### Reading the Configuration

When the UI initializes, the controller reads the `asr_engine` key from `app/config-user.json5`. If the key is absent, the system defaults to `'faster-whisper'`:

```python
asr_engine = self.user_config.get("asr_engine", 'faster-whisper')
self.whisper_inf = self.switch_case(asr_engine)

```

This logic appears in both [`app/gradio_gulliver.py`](https://github.com/abus-aikorea/voice-pro/blob/main/app/gradio_gulliver.py) (lines 59-61) and [`app/gradio_asr.py`](https://github.com/abus-aikorea/voice-pro/blob/main/app/gradio_asr.py) (lines 29-31), ensuring consistent behavior across the full dubbing pipeline and the dedicated ASR tab.

### Persisting User Preferences

After a user selects an engine and model, the controller writes the values back to the configuration:

```python
self.user_config.set("asr_engine", asr_engine)
self.user_config.set(f'{asr_engine.replace("-", "_")}_model', modelName)

```

This persistence mechanism, found in [`app/gradio_gulliver.py`](https://github.com/abus-aikorea/voice-pro/blob/main/app/gradio_gulliver.py) (lines 71-74), ensures subsequent sessions initialize with the previously selected ASR engine.

## The Factory Pattern Implementation

Voice-Pro decouples engine selection from instantiation using a private factory method that maps string identifiers to concrete inference classes.

### The switch_case Method

Both `GradioGulliver` and `GradioASR` implement a `switch_case` method that acts as a factory:

```python
def switch_case(self, case):
    switch_dict = {
        'faster-whisper':   lambda: FasterWhisperInference(),
        'whisper':          lambda: WhisperInference(),
        'whisper-timestamped': lambda: WhisperTimestampedInference()
    }
    return switch_dict.get(case, lambda: FasterWhisperInference())()

```

This implementation appears in [`app/gradio_gulliver.py`](https://github.com/abus-aikorea/voice-pro/blob/main/app/gradio_gulliver.py) (lines 67-73) and [`app/gradio_asr.py`](https://github.com/abus-aikorea/voice-pro/blob/main/app/gradio_asr.py) (lines 36-42). The method returns a lambda instantiation of the appropriate wrapper class, defaulting to `FasterWhisperInference` if an invalid identifier is provided.

## ASR Engine Implementations

Voice-Pro provides three specialized wrapper classes that expose a unified interface for model discovery and transcription.

### Faster-Whisper Wrapper

The [`app/abus_asr_faster_whisper.py`](https://github.com/abus-aikorea/voice-pro/blob/main/app/abus_asr_faster_whisper.py) module wraps the Faster-Whisper model, offering optimized inference with methods such as `available_models()`, `available_langs()`, and `transcribe_file()`.

### Original Whisper Wrapper

Located in [`app/abus_asr_whisper.py`](https://github.com/abus-aikorea/voice-pro/blob/main/app/abus_asr_whisper.py), this wrapper provides a thin interface to the original OpenAI Whisper model, maintaining the same method signatures for drop-in compatibility.

### Whisper-Timestamped Wrapper

The [`app/abus_asr_whisper_timestamped.py`](https://github.com/abus-aikorea/voice-pro/blob/main/app/abus_asr_whisper_timestamped.py) module extends the base Whisper functionality to provide word-level timestamp alignment, essential for precise subtitle generation.

### Unified API Interface

All three implementations expose a consistent interface:

- `available_models()` – Returns a list of supported model sizes
- `available_langs()` – Returns supported language codes  
- `transcribe_file()` – Executes speech-to-text conversion

This uniform API allows the Gradio controllers to treat any engine interchangeably without conditional logic scattered throughout the codebase.

## Runtime Transcription Workflow

When a user initiates transcription, the controller instantiates the selected engine and executes the inference pipeline.

### Engine Instantiation

The controller re-instantiates the inference class whenever the engine selection changes:

```python
self.whisper_inf = self.switch_case(asr_engine)

```

This ensures the correct backend is loaded with the appropriate model weights and compute settings.

### Executing Transcription

The transcription workflow in [`app/gradio_gulliver.py`](https://github.com/abus-aikorea/voice-pro/blob/main/app/gradio_gulliver.py) (lines 85-91) demonstrates the unified API in action:

```python
self.whisper_inf = self.switch_case(asr_engine)
subtitles = self.whisper_inf.transcribe_file(input_path, params, False, gr.Progress())

```

The `transcribe_file` method accepts the audio path, parameters (model size, language, compute type), a boolean flag, and a Gradio progress tracker.

### Complete Usage Example

To run transcription programmatically with a specific engine:

```python
params = WhisperParameters(
    model_size="large",
    lang="english",
    compute_type="float16"
)
audio_path = "/workspace/audio.wav"
subtitles = self.whisper_inf.transcribe_file(audio_path, params, False, gr.Progress())

```

## Dynamic UI Updates

The interface adapts dynamically when users switch engines, updating available models and languages accordingly.

### Refreshing Model Lists

When the engine selector changes, the controller calls `update_whisper_models()` to repopulate the model dropdown:

```python
def update_whisper_models(self, asr_engine):
    whisper_inf = self.switch_case(asr_engine)
    model_list = whisper_inf.available_models()
    # Returns Gradio update object with new choices

```

This method, found in [`app/gradio_gulliver.py`](https://github.com/abus-aikorea/voice-pro/blob/main/app/gradio_gulliver.py) (lines 86-93), creates a temporary instance of the selected engine to query its `available_models()` method, ensuring the UI only presents compatible options.

## Summary

- Voice-Pro supports three ASR engines: **Faster-Whisper**, **Whisper**, and **Whisper-Timestamped**, configurable via `app/config-user.json5`.
- The `switch_case` factory method in `GradioGulliver` and `GradioASR` maps string identifiers to concrete inference classes, defaulting to Faster-Whisper.
- Each engine wrapper ([`app/abus_asr_faster_whisper.py`](https://github.com/abus-aikorea/voice-pro/blob/main/app/abus_asr_faster_whisper.py), [`app/abus_asr_whisper.py`](https://github.com/abus-aikorea/voice-pro/blob/main/app/abus_asr_whisper.py), [`app/abus_asr_whisper_timestamped.py`](https://github.com/abus-aikorea/voice-pro/blob/main/app/abus_asr_whisper_timestamped.py)) implements a unified API for model discovery and transcription.
- Runtime transcription uses the `transcribe_file()` method with consistent parameters regardless of the selected backend.
- User selections persist across sessions through the `UserConfig` class, and the UI dynamically updates model lists based on the active engine.

## Frequently Asked Questions

### How do I switch ASR engines in Voice-Pro?

Navigate to the engine dropdown in the Gradio interface and select your preferred backend. The controller stores this choice in `app/config-user.json5` under the `asr_engine` key, then reinstantiates the inference class via the `switch_case` factory method.

### What is the default ASR engine in Voice-Pro?

The system defaults to **Faster-Whisper** if no previous selection exists. This default is hardcoded in the `user_config.get("asr_engine", 'faster-whisper')` call found in both [`app/gradio_gulliver.py`](https://github.com/abus-aikorea/voice-pro/blob/main/app/gradio_gulliver.py) and [`app/gradio_asr.py`](https://github.com/abus-aikorea/voice-pro/blob/main/app/gradio_asr.py).

### Can I use Whisper-Timestamped for word-level alignment?

Yes. Select **Whisper-Timestamped** from the engine dropdown. The [`app/abus_asr_whisper_timestamped.py`](https://github.com/abus-aikorea/voice-pro/blob/main/app/abus_asr_whisper_timestamped.py) wrapper extends the base Whisper functionality to provide precise timestamped output suitable for subtitle generation.

### Where is the ASR engine selection logic implemented?

The core selection logic resides in [`app/gradio_gulliver.py`](https://github.com/abus-aikorea/voice-pro/blob/main/app/gradio_gulliver.py) (for the full dubbing pipeline) and [`app/gradio_asr.py`](https://github.com/abus-aikorea/voice-pro/blob/main/app/gradio_asr.py) (for the ASR-only tab). Both files contain the `switch_case` factory method and configuration persistence code.