Understanding the TranscriptionEngine Protocol Architecture in OpenSuperWhisper
The TranscriptionEngine protocol is a minimal Swift abstraction that decouples transcription workflows from specific implementations by defining five core requirements: model state tracking, async initialization, audio transcription, cancellation support, and language enumeration.
The OpenSuperWhisper macOS application leverages protocol-oriented design to remain agnostic of underlying speech-to-text providers. At the center of this architecture sits the TranscriptionEngine protocol, which standardizes how the UI layer interacts with complex inference engines without coupling to specific technologies like Whisper or Silero VAD. This design enables seamless swapping between local and cloud-based transcription backends through simple dependency injection.
Core Protocol Definition
The protocol is defined in OpenSuperWhisper/Engines/TranscriptionEngine.swift as a minimal interface that any transcription backend must implement:
protocol TranscriptionEngine: AnyObject {
var isModelLoaded: Bool { get }
var engineName: String { get }
func initialize() async throws
func transcribeAudio(url: URL, settings: Settings) async throws -> String
func cancelTranscription()
func getSupportedLanguages() -> [String]
}
This interface deliberately excludes implementation details such as model formats, audio codecs, or inference frameworks. By conforming to these five members, any class—from local Whisper instances to cloud API clients—can plug into the existing transcription service without modifying higher-level code.
Architectural Responsibilities
The protocol divides transcription workflows into five distinct responsibilities, each implemented differently by concrete engines.
Model Lifecycle Management
The isModelLoaded property and initialize() method create a two-phase startup pattern. Callers can query whether a model is resident in memory before attempting transcription, while initialize() handles expensive setup operations asynchronously. In WhisperEngine.swift, the implementation resolves the model path from AppPreferences, instantiates a MyWhisperContext, and validates that context != nil before marking the engine ready.
Audio Transcription Pipeline
The transcribeAudio(url:settings:) method serves as the primary entry point for end-to-end processing. The method signature accepts a file URL and a Settings configuration object, returning a String containing the final transcription. Concrete implementations handle the entire pipeline internally, including format conversion, voice activity detection, and post-processing. This encapsulation allows the UI layer to remain unaware of PCM conversion, VAD segmentation, or beam search parameters.
Cancellation Support
The cancelTranscription() method provides cooperative cancellation for long-running inference tasks. Implementations must guarantee that calling this method promptly terminates ongoing work. WhisperEngine achieves this through an AbortFlag class—a thread-safe wrapper around an NSLock-protected Boolean that the Whisper C API checks via an abortCallback. When flipped, the native loop aborts and returns control to Swift.
Language Enumeration
getSupportedLanguages() returns an array of language codes the engine can process. This enables UI features like language selection dropdowns without hardcoding supported locales. WhisperEngine delegates this call to LanguageUtil.availableLanguages, centralizing the language list in LanguageUtil.swift.
Extensibility Design
The protocol's minimal surface area maximizes extensibility. Adding a streaming-oriented engine like FluidAudioEngine requires implementing only these five members, without modifying TranscriptionService.swift or view controllers. This plug-and-play architecture supports future backends including cloud APIs or specialized medical transcription models.
WhisperEngine Implementation Details
The concrete WhisperEngine class in OpenSuperWhisper/Engines/WhisperEngine.swift demonstrates how a complex local inference pipeline conforms to the simple protocol interface.
Audio Preparation and VAD
Before inference, convertAudioToPCM(_:) transforms input files into 16 kHz mono PCM format, normalizing multi-channel audio and exploiting parallel conversion for long files. The engine then runs detectSpeech(in:) using the Silero VAD model (ggml-silero-v5.1.2.bin) to extract speech-only segments. The helper speechOnlySamples(from:segments:) stitches these segments while preserving natural pauses, reducing Whisper's processing load and improving accuracy.
Progress Reporting
A ProgressContext struct bridges the C-based Whisper callbacks with Swift concurrency. The Whisper API invokes progressCallback, which maps the native 0-100% range to a UI-friendly 10-95% range and dispatches updates via DispatchQueue.main.async. This keeps the protocol interface clean while still providing granular feedback.
Post-Processing Pipeline
After inference, WhisperEngine applies several normalization steps: stripping placeholder tags like [MUSIC] and [BLANK_AUDIO], trimming whitespace, and optionally running the Asian autocorrection wrapper defined in AutocorrectWrapper.swift. These transformations happen entirely within the concrete implementation, ensuring the protocol method always returns clean, user-ready text.
Integration Examples
Basic Transcription Workflow
The following pattern demonstrates dependency injection using the protocol:
import OpenSuperWhisper
func transcribeFile(at url: URL) async {
let engine: TranscriptionEngine = WhisperEngine()
do {
try await engine.initialize()
let settings = Settings(
selectedLanguage: "en",
useBeamSearch: false,
temperature: 0.0,
showTimestamps: true,
suppressBlankAudio: false,
initialPrompt: "",
beamSize: 5,
noSpeechThreshold: 0.6,
shouldApplyAsianAutocorrect: true
)
let text = try await engine.transcribeAudio(url: url, settings: settings)
print("Transcription result:\n\(text)")
} catch {
print("Transcription failed: \(error)")
}
}
Implementing Cancellation
To cancel an ongoing transcription, call the protocol method from any concurrency context:
let engine: TranscriptionEngine = WhisperEngine()
Task {
do {
try await engine.initialize()
let transcriptionTask = Task {
try await engine.transcribeAudio(url: audioURL, settings: settings)
}
try await Task.sleep(nanoseconds: 2_000_000_000)
engine.cancelTranscription()
let _ = await transcriptionTask.result
} catch {
print("Cancelled or failed: \(error)")
}
}
Querying Language Support
Access supported languages without instantiating model resources:
let engine: TranscriptionEngine = WhisperEngine()
let languages = engine.getSupportedLanguages()
print("Available languages: \(languages)")
Key Source Files
TranscriptionEngine.swift– Defines the core protocol abstraction that all engines must implement.WhisperEngine.swift– Concrete implementation handling Whisper model loading, VAD, and audio conversion.FluidAudioEngine.swift– Alternative streaming-oriented engine demonstrating protocol extensibility.Settings.swift– Configuration container for transcription parameters passed totranscribeAudio.LanguageUtil.swift– Utility class providing the language list forgetSupportedLanguages().AutocorrectWrapper.swift– Post-processing helper for Asian language corrections.
Summary
- The TranscriptionEngine protocol decouples transcription workflows from specific implementations through a five-member interface.
- Model lifecycle methods (
initialize(),isModelLoaded) enable expensive setup without blocking UI threads. - Cancellation is implemented via thread-safe flags (
AbortFlag) that bridge Swift concurrency with C callbacks. - Audio processing pipelines (conversion, VAD, inference) remain encapsulated behind the simple
transcribeAudio(url:settings:)signature. - Extensibility is achieved through minimal protocol requirements, allowing new engines like
FluidAudioEngineto integrate without architectural changes.
Frequently Asked Questions
What are the required methods and properties in the TranscriptionEngine protocol?
The protocol requires two properties (isModelLoaded: Bool, engineName: String) and three methods (initialize(), transcribeAudio(url:settings:), cancelTranscription(), getSupportedLanguages()). Any class conforming to these five members can function as a transcription backend for the application.
How does WhisperEngine handle cancellation during transcription?
WhisperEngine implements cancelTranscription() by setting an AbortFlag, a thread-safe Boolean protected by NSLock. The flag is passed to Whisper's C API via an abortCallback pointer; when the flag is set, the native inference loop terminates early and returns control to Swift with a cancellation error.
Can I implement a custom transcription engine for OpenSuperWhisper?
Yes. Create a new class conforming to TranscriptionEngine and implement the five required members. The protocol's minimal design allows integration of cloud APIs, specialized medical models, or real-time streaming engines without modifying TranscriptionService.swift or UI code.
What is the difference between TranscriptionEngine and WhisperEngine?
TranscriptionEngine is the protocol abstraction defining what capabilities a transcription provider must offer. WhisperEngine is the concrete implementation that satisfies how those capabilities work using local Whisper models, Silero VAD, and PCM audio conversion. This separation allows the UI to depend on the protocol while swapping concrete implementations at runtime.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →