Where Is the Whisper Engine Implemented in OpenSuperWhisper?
The Whisper engine in OpenSuperWhisper is implemented in the OpenSuperWhisper/Engines package, specifically within the WhisperEngine.swift file, which orchestrates model loading, audio conversion, and inference by wrapping the native C library via MyWhisperContext found in Whis.swift.
OpenSuperWhisper is a macOS application that provides real-time speech recognition using OpenAI’s Whisper models. To understand exactly where the Whisper engine is implemented in OpenSuperWhisper, developers must examine the Swift-layer abstraction that bridges the UI with the underlying C++ inference code. The implementation spans two primary files that handle high-level orchestration and low-level library binding.
Core Engine Architecture
The transcription capability is split between a high-level manager and a thin native wrapper. This separation allows the UI to interact with a clean Swift API while the heavy lifting occurs in the optimized C++ Whisper implementation.
The WhisperEngine Class
The primary entry point for all transcription operations is the WhisperEngine class, defined in OpenSuperWhisper/Engines/WhisperEngine.swift. This class conforms to the TranscriptionEngine protocol, ensuring a consistent interface for different engine types. According to the OpenSuperWhisper source code, WhisperEngine handles six critical responsibilities:
- Model initialization – The
initialize()method (lines 66‑73) reads the selected.binmodel file and instantiates aMyWhisperContextobject. - Audio preprocessing – The
convertAudioToPCM(_:)method (lines 39‑84) converts any input audio into 16 kHz mono Float‑32 PCM, the exact format the Whisper C API expects. - Parameter configuration – A
WhisperFullParamsstruct is populated from UISettings(lines 16‑28), configuring language, beam search width, temperature, and other inference hyperparameters. - Progress propagation – Custom C‑compatible callbacks forward Whisper’s internal progress (0‑100 %) to Swift via a
ProgressContextobject (lines 29‑55), enabling real‑time UI updates. - Inference execution – The engine calls
context.full(samples:params:)(lines 71‑74) to run the actual neural network transcription on the PCM buffer. - Result processing – Post‑inference, the engine iterates over segments, optionally injects timestamps, cleans up control markers, and applies Asian‑language autocorrection (lines 78‑107).
The Low-Level Bridge: MyWhisperContext
Underlying WhisperEngine is MyWhisperContext, located in OpenSuperWhisper/Whis/Whis.swift. This thin Swift class wraps the native Whisper C API, exposing methods such as initFromFile, full, fullNSegments, and segment getters. Lines 18‑45 of Whis.swift define these bindings, translating Swift data types into the C structures required by the underlying whisper.cpp implementation. Together, these two files constitute the complete Whisper engine implementation within OpenSuperWhisper.
Step-by-Step Transcription Flow
When a user initiates transcription, the engine executes a deterministic pipeline:
- Service Initialization –
TranscriptionService.swiftcreates and loads the selected engine, wiring UI progress callbacks to the engine’sonProgressUpdatehandler. - Model Loading –
WhisperEngine.initialize()validates the model file path and constructsMyWhisperContextviainitFromFile. - Format Conversion – Input audio passes through
convertAudioToPCM(_: ), which resamples to 16 kHz and converts to Float‑32 PCM using AVFoundation. - Parameter Binding – UI settings map directly to
WhisperFullParamsfields, including language detection and decoding strategies. - Native Inference – The PCM buffer and parameters pass to
context.full(samples:params:), which blocks until the C++ library completes inference. - Segment Assembly – The engine queries
fullNSegmentsand iterates through each segment via getter methods, concatenating text and applying post‑processing.
Key Source Files and Responsibilities
Understanding the file structure clarifies where each piece of the engine lives:
OpenSuperWhisper/Engines/WhisperEngine.swift– High‑level Swift engine implementing theTranscriptionEngineprotocol; manages model lifecycle, audio conversion, progress callbacks, and result assembly.OpenSuperWhisper/Whis/Whis.swift– Thin wrapper around the native Whisper C library (MyWhisperContext); exposes initialization, full transcription, and segment access methods.OpenSuperWhisper/Engines/TranscriptionEngine.swift– Protocol definition establishing the common interface forWhisperEngineand alternative implementations likeFluidAudioEngine.OpenSuperWhisper/TranscriptionService.swift– Application‑level service responsible for engine instantiation, lifecycle management, and binding UI events to engine callbacks.
Practical Usage Example
The following Swift code demonstrates how to instantiate the engine, configure settings, and execute transcription with progress tracking:
import Foundation
// Initialize the engine (typically handled by TranscriptionService)
let engine = WhisperEngine()
await engine.initialize()
// Configure UI feedback
engine.onProgressUpdate = { progress in
print("Transcription progress: \(Int(progress * 100))%")
}
// Execute transcription
let audioURL = URL(fileURLWithPath: "/path/to/recording.wav")
let settings = Settings() // Language, beam size, etc.
let result = try await engine.transcribeAudio(url: audioURL, settings: settings)
print("Transcribed text: \(result)")
Summary
- The core Whisper engine is implemented in
OpenSuperWhisper/Engines/WhisperEngine.swift. - Low-level C API bindings reside in
OpenSuperWhisper/Whis/Whis.swiftwithin theMyWhisperContextclass. - The engine conforms to the
TranscriptionEngineprotocol, enabling pluggable architecture. - Audio conversion, parameter preparation, and progress callbacks are all handled within
WhisperEnginebefore delegating inference to the native library. - Result assembly includes timestamp formatting and language‑specific post‑processing.
Frequently Asked Questions
Where is the Whisper model file loaded in OpenSuperWhisper?
The model file is loaded inside WhisperEngine.swift at lines 66‑73 via the initialize() method, which constructs MyWhisperContext by calling initFromFile with the path to the selected .bin model.
How does OpenSuperWhisper convert audio for Whisper compatibility?
The convertAudioToPCM(_:) method in WhisperEngine.swift (lines 39‑84) handles all preprocessing, using AVFoundation to resample input audio to 16 kHz mono and convert it to Float‑32 PCM format required by the Whisper C API.
What is the relationship between WhisperEngine and MyWhisperContext?
WhisperEngine is the high-level Swift coordinator that manages UI interactions, settings, and audio preprocessing, while MyWhisperContext is the low-level bridge in Whis.swift that directly wraps the native Whisper C library functions for model initialization and inference.
How can I track transcription progress in OpenSuperWhisper?
Assign a closure to the onProgressUpdate property of WhisperEngine. The engine forwards progress from the C library (0.0 to 1.0) through a ProgressContext object, allowing real-time updates to UI elements during the context.full(samples:params:) call.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →