Where Is the Whisper Engine Implemented in OpenSuperWhisper?

The Whisper engine in OpenSuperWhisper is implemented in the OpenSuperWhisper/Engines package, specifically within the WhisperEngine.swift file, which orchestrates model loading, audio conversion, and inference by wrapping the native C library via MyWhisperContext found in Whis.swift.

OpenSuperWhisper is a macOS application that provides real-time speech recognition using OpenAI’s Whisper models. To understand exactly where the Whisper engine is implemented in OpenSuperWhisper, developers must examine the Swift-layer abstraction that bridges the UI with the underlying C++ inference code. The implementation spans two primary files that handle high-level orchestration and low-level library binding.

Core Engine Architecture

The transcription capability is split between a high-level manager and a thin native wrapper. This separation allows the UI to interact with a clean Swift API while the heavy lifting occurs in the optimized C++ Whisper implementation.

The WhisperEngine Class

The primary entry point for all transcription operations is the WhisperEngine class, defined in OpenSuperWhisper/Engines/WhisperEngine.swift. This class conforms to the TranscriptionEngine protocol, ensuring a consistent interface for different engine types. According to the OpenSuperWhisper source code, WhisperEngine handles six critical responsibilities:

  • Model initialization – The initialize() method (lines 66‑73) reads the selected .bin model file and instantiates a MyWhisperContext object.
  • Audio preprocessing – The convertAudioToPCM(_:) method (lines 39‑84) converts any input audio into 16 kHz mono Float‑32 PCM, the exact format the Whisper C API expects.
  • Parameter configuration – A WhisperFullParams struct is populated from UI Settings (lines 16‑28), configuring language, beam search width, temperature, and other inference hyperparameters.
  • Progress propagation – Custom C‑compatible callbacks forward Whisper’s internal progress (0‑100 %) to Swift via a ProgressContext object (lines 29‑55), enabling real‑time UI updates.
  • Inference execution – The engine calls context.full(samples:params:) (lines 71‑74) to run the actual neural network transcription on the PCM buffer.
  • Result processing – Post‑inference, the engine iterates over segments, optionally injects timestamps, cleans up control markers, and applies Asian‑language autocorrection (lines 78‑107).

The Low-Level Bridge: MyWhisperContext

Underlying WhisperEngine is MyWhisperContext, located in OpenSuperWhisper/Whis/Whis.swift. This thin Swift class wraps the native Whisper C API, exposing methods such as initFromFile, full, fullNSegments, and segment getters. Lines 18‑45 of Whis.swift define these bindings, translating Swift data types into the C structures required by the underlying whisper.cpp implementation. Together, these two files constitute the complete Whisper engine implementation within OpenSuperWhisper.

Step-by-Step Transcription Flow

When a user initiates transcription, the engine executes a deterministic pipeline:

  1. Service Initialization – TranscriptionService.swift creates and loads the selected engine, wiring UI progress callbacks to the engine’s onProgressUpdate handler.
  2. Model Loading – WhisperEngine.initialize() validates the model file path and constructs MyWhisperContext via initFromFile.
  3. Format Conversion – Input audio passes through convertAudioToPCM(_: ), which resamples to 16 kHz and converts to Float‑32 PCM using AVFoundation.
  4. Parameter Binding – UI settings map directly to WhisperFullParams fields, including language detection and decoding strategies.
  5. Native Inference – The PCM buffer and parameters pass to context.full(samples:params:), which blocks until the C++ library completes inference.
  6. Segment Assembly – The engine queries fullNSegments and iterates through each segment via getter methods, concatenating text and applying post‑processing.

Key Source Files and Responsibilities

Understanding the file structure clarifies where each piece of the engine lives:

Practical Usage Example

The following Swift code demonstrates how to instantiate the engine, configure settings, and execute transcription with progress tracking:

import Foundation

// Initialize the engine (typically handled by TranscriptionService)
let engine = WhisperEngine()
await engine.initialize()

// Configure UI feedback
engine.onProgressUpdate = { progress in
    print("Transcription progress: \(Int(progress * 100))%")
}

// Execute transcription
let audioURL = URL(fileURLWithPath: "/path/to/recording.wav")
let settings = Settings() // Language, beam size, etc.
let result = try await engine.transcribeAudio(url: audioURL, settings: settings)
print("Transcribed text: \(result)")

Summary

  • The core Whisper engine is implemented in OpenSuperWhisper/Engines/WhisperEngine.swift.
  • Low-level C API bindings reside in OpenSuperWhisper/Whis/Whis.swift within the MyWhisperContext class.
  • The engine conforms to the TranscriptionEngine protocol, enabling pluggable architecture.
  • Audio conversion, parameter preparation, and progress callbacks are all handled within WhisperEngine before delegating inference to the native library.
  • Result assembly includes timestamp formatting and language‑specific post‑processing.

Frequently Asked Questions

Where is the Whisper model file loaded in OpenSuperWhisper?

The model file is loaded inside WhisperEngine.swift at lines 66‑73 via the initialize() method, which constructs MyWhisperContext by calling initFromFile with the path to the selected .bin model.

How does OpenSuperWhisper convert audio for Whisper compatibility?

The convertAudioToPCM(_:) method in WhisperEngine.swift (lines 39‑84) handles all preprocessing, using AVFoundation to resample input audio to 16 kHz mono and convert it to Float‑32 PCM format required by the Whisper C API.

What is the relationship between WhisperEngine and MyWhisperContext?

WhisperEngine is the high-level Swift coordinator that manages UI interactions, settings, and audio preprocessing, while MyWhisperContext is the low-level bridge in Whis.swift that directly wraps the native Whisper C library functions for model initialization and inference.

How can I track transcription progress in OpenSuperWhisper?

Assign a closure to the onProgressUpdate property of WhisperEngine. The engine forwards progress from the C library (0.0 to 1.0) through a ProgressContext object, allowing real-time updates to UI elements during the context.full(samples:params:) call.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →