What Speech Recognition Models Does Voice-Pro Support? A Complete Guide to Whisper Back-Ends

Voice-Pro supports three interchangeable ASR back-ends—faster-whisper, whisper, and whisper-timestamped—each offering the full range of official Whisper model sizes from tiny to large-v3.

Voice-Pro is an open-source audio processing toolkit that provides flexible speech recognition through multiple Whisper-based engines. The repository (abus-aikorea/voice-pro) ships with three distinct ASR back-ends that users can swap via the Gradio interface or programmatically through Python. Understanding which speech recognition models are available and how to configure them is essential for optimizing transcription speed, accuracy, and memory usage.

The Three ASR Engines in Voice-Pro

Voice-Pro implements a pluggable architecture where you select your preferred speech recognition engine through the asr_engine_radio UI component defined in app/gradio_live_translate.py (lines 28-40) or via configuration files. Each engine wraps a specific Whisper implementation while exposing a unified interface through the WhisperParameters class defined in app/abus_asr_parameters.py.

faster-whisper (FasterWhisperInference)

The faster-whisper back-end, implemented in app/abus_asr_faster_whisper.py, utilizes CTranslate2-optimized inference for significantly faster processing. The FasterWhisperInference class provides available_models() which returns the official Whisper model families: "tiny", "base", "small", "medium", "large", and newer variants like "large-v1", "large-v2", "large-v3" (lines 84-86). This engine also exposes available_compute_types() (lines 92-112) supporting quantization options such as float16 and int8_float16 to reduce VRAM usage.

whisper (WhisperInference)

The original OpenAI reference implementation is available through WhisperInference in app/abus_asr_whisper.py (lines 43-51). This back-end calls whisper.available_models() and provides the same standard model sizes as the faster variant but without CTranslate2 optimization. It serves as the baseline implementation for users requiring strict compatibility with the original Whisper behavior, with model listing implemented in WhisperInference.available_models() (lines 71-74).

whisper-timestamped (WhisperTimestampedInference)

For applications requiring precise word-level alignment, the whisper-timestamped engine in app/abus_asr_whisper_timestamped.py extends the base functionality with granular timestamp output. The WhisperTimestampedInference class adds per-word timestamps to the transcription segments while supporting the identical model catalog as the other two engines.

Available Whisper Model Sizes

All three engines support the complete Whisper model family, allowing you to trade accuracy for speed based on your hardware constraints. The models range from tiny (39M parameters, fastest) to large-v3 (1.5B parameters, most accurate). You can retrieve the available list programmatically using the available_models() class method on any inference class.

Configuration and Common Parameters

All engines share the WhisperParameters configuration interface from app/abus_asr_parameters.py. This dataclass standardizes parameters including model_size, compute_type (faster-whisper only), lang, is_translate, and denoise_level. The UI populates dropdown menus using helper methods like available_langs() (defined in both FasterWhisperInference lines 87-89 and WhisperInference lines 75-78) to display supported languages (e.g., "English", "Korean") and available_compute_types() for quantization settings.

Practical Code Examples

To instantiate a specific ASR engine and list supported models:

from app.abus_asr_faster_whisper import FasterWhisperInference
from app.abus_asr_whisper import WhisperInference

print("faster-whisper models:", FasterWhisperInference.available_models())
print("whisper models:", WhisperInference.available_models())

To transcribe audio using the faster-whisper engine with specific compute settings:

from app.abus_asr_faster_whisper import FasterWhisperInference
from app.abus_asr_parameters import WhisperParameters

asr = FasterWhisperInference()
params = WhisperParameters(
    model_size="large-v3",
    compute_type="float16",
    lang="English",
    is_translate=False,
    denoise_level=0,
)
segments, duration = asr.transcribe_file("sample.wav", params, highlight_words=True)
print(f"Transcribed {len(segments)} segments in {duration:.2f}s")

For word-level timestamp alignment using the timestamped engine:

from app.abus_asr_whisper_timestamped import WhisperTimestampedInference
from app.abus_asr_parameters import WhisperParameters

asr_ts = WhisperTimestampedInference()
params = WhisperParameters(
    model_size="medium",
    lang="Korean",
    is_translate=False,
    denoise_level=0,
)
segments, _ = asr_ts.transcribe_file("korean_clip.wav", params, highlight_words=False)

# Each segment contains 'words' with precise start/end timestamps

Summary

  • Voice-Pro provides three ASR back-ends: faster-whisper, whisper, and whisper-timestamped, selectable via the asr_engine_radio UI control in app/gradio_live_translate.py or programmatically.
  • All engines support the full range of official Whisper model sizes from tiny to large-v3 through their respective available_models() methods.
  • The faster-whisper engine offers additional compute type quantization (float16, int8_float16) for hardware-constrained environments via FasterWhisperInference.available_compute_types().
  • Configuration is unified through the WhisperParameters class in app/abus_asr_parameters.py, ensuring consistent API access across all back-ends.
  • Word-level timestamps are exclusively available through the whisper-timestamped engine implemented in app/abus_asr_whisper_timestamped.py.

Frequently Asked Questions

What is the difference between faster-whisper and whisper in Voice-Pro?

The faster-whisper engine uses CTranslate2 optimization in app/abus_asr_faster_whisper.py for significantly faster inference and supports quantization via available_compute_types(), while the whisper engine in app/abus_asr_whisper.py runs the original OpenAI implementation without these optimizations but maintains strict reference compatibility.

Which Whisper model size should I choose for my hardware?

Choose tiny or base for CPU-only environments or real-time applications requiring low latency. Select medium or large variants when accuracy is critical and GPU memory is available (8GB+ VRAM recommended for large-v3). Use available_compute_types() with faster-whisper to enable int8 quantization if VRAM is limited.

How do I enable word-level timestamps in Voice-Pro?

Use the whisper-timestamped engine by selecting it in the UI or instantiating WhisperTimestampedInference from app/abus_asr_whisper_timestamped.py. When calling transcribe_file(), the returned segments will contain nested words arrays with precise millisecond-level timestamps for each token.

Does Voice-Pro support quantized models for faster inference?

Yes, but only when using the faster-whisper engine. The FasterWhisperInference.available_compute_types() method returns supported CTranslate2 quantization options including float16, int8, and int8_float16, which you can pass via WhisperParameters.compute_type to reduce memory footprint and increase throughput.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →