ASR Engines Supported by VoiceStudio: Complete Registry Guide

VoiceStudio supports 10 distinct ASR engines including WhisperX (default), Faster-Whisper, MLX Whisper, PyTorch Whisper, NVIDIA NeMo Parakeet, Moonshine, FunASR, sherpa-onnx, and OpenAI-compatible endpoints.

VoiceStudio is an open-source audio processing platform that abstracts multiple Automatic Speech Recognition backends behind a unified interface. Understanding which ASR engines are supported by VoiceStudio enables developers to select optimal implementations for offline transcription, hardware-accelerated inference, or low-latency live dictation.

Complete List of ASR Engines Supported by VoiceStudio

The canonical registry resides in docs/features.yaml (lines 64-86), which serves as the source of truth for documentation and UI generation. At runtime, these engines are discovered via the _REGISTRY dictionary in backend/services/asr_backend.py.

VoiceStudio supports the following engine IDs:

  • whisperx — WhisperX (default): The primary offline ASR engine used for standard speech-to-text tasks.
  • faster-whisper — Faster-Whisper: An optimized Whisper variant engineered for faster inference speeds.
  • mlx-whisper — MLX Whisper: Implementation utilizing Apple's MLX framework for hardware acceleration on Apple Silicon.
  • pytorch-whisper — PyTorch Whisper: The standard PyTorch-based Whisper model implementation.
  • nemo-parakeet — Parakeet TDT: NVIDIA NeMo-based ASR delivering high-accuracy transcription.
  • parakeet-mlx — Parakeet TDT v3 (MLX): MLX-accelerated variant of the Parakeet model for Apple hardware.
  • moonshine — Moonshine: Lightweight ASR model designed for resource-constrained environments.
  • funasr — FunASR: Open-source ASR with integrated speaker diarization support.
  • sherpa-onnx-asr — sherpa-onnx: ONNX-based engine specifically optimized for low-latency live dictation.
  • openai-compat-asr — OpenAI-compatible: Connector for external OpenAI-compatible ASR services requiring user configuration.

Each engine ID maps to a concrete implementation class conforming to the ASRBackend interface, allowing the frontend to present a single unified "ASR" model entry regardless of the underlying provider.

ASR Engine Registration Architecture

VoiceStudio implements a registry pattern to manage its supported ASR engines. The backend/services/asr_backend.py file defines a _REGISTRY dictionary that maps string identifiers to backend classes.

During initialization, the application reads the asr_engines section from docs/features.yaml to populate the user interface. This manifest-based approach decouples feature documentation from implementation code. The ASRBackend abstract interface standardizes initialization, inference, and device management methods across all implementations, ensuring consistent behavior whether using local WhisperX or remote OpenAI-compatible endpoints.

The UI localization strings in frontend/src/i18n/locales/en.json define "role_asr": "ASR" for the model panel, while backend/services/model_lifecycle.py adds specific entries like "WhisperX ASR" to the loaded-models panel with device and checkpoint information.

Accessing ASR Engines Programmatically

Developers can query available engines and initialize specific backends at runtime using the public API exposed in backend/services/asr_backend.py.

List all supported engine IDs:

from backend.services import asr_backend

# The registry maps engine_id to BackendClass

available_asr_ids = list(asr_backend._REGISTRY.keys())
print("Supported ASR engines:", available_asr_ids)

Initialize a specific backend:

from backend.services import asr_backend

engine_id = "funasr"
BackendClass = asr_backend._REGISTRY[engine_id]
backend = BackendClass()

print(f"Initialized {engine_id}")

Refer to the ASRBackend class definition for specific public methods regarding model loading and transcription.

Summary

VoiceStudio provides comprehensive ASR integration through its registry-based architecture:

  • Ten supported engines ranging from local Whisper variants to remote OpenAI-compatible endpoints
  • Manifest-driven configuration maintained in docs/features.yaml (lines 64-86)
  • Runtime discovery via _REGISTRY in backend/services/asr_backend.py
  • Unified abstraction through the ASRBackend interface contract
  • Hardware-specific optimizations including MLX support for Apple Silicon and ONNX for low-latency dictation

Frequently Asked Questions

What is the default ASR engine in VoiceStudio?

WhisperX (whisperx) serves as the default offline ASR engine, automatically selected for standard speech-to-text tasks according to the feature manifest in docs/features.yaml.

How does VoiceStudio switch between ASR engines?

The UI model selection panel reads from the _REGISTRY dictionary in backend/services/asr_backend.py. Each engine appears as a unified "ASR" entry in the interface, with the specific implementation class instantiated based on the selected engine ID.

Can VoiceStudio connect to external ASR services?

Yes. The openai-compat-asr engine ID enables connections to external OpenAI-compatible ASR endpoints. Configure your server URL and authentication details in the application settings to route transcription requests remotely.

Use sherpa-onnx (sherpa-onnx-asr). This ONNX-based engine is specifically optimized for low-latency live dictation scenarios, providing significantly faster response times than offline-focused alternatives like standard WhisperX.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →