OpenSuperWhisper WhisperEngine.swift: Complete Guide to Speech-to-Text Implementation

WhisperEngine.swift implements the concrete Whisper transcription engine that powers OpenSuperWhisper's speech-to-text workflow, providing model loading, 16kHz mono PCM audio conversion, real-time progress reporting, and thread-safe cancellation through a protocol-based architecture.

OpenSuperWhisper leverages WhisperEngine.swift as its primary transcription backend, conforming to the TranscriptionEngine protocol to deliver a complete pipeline from raw audio files to processed text output. Located at OpenSuperWhisper/Engines/WhisperEngine.swift, this file bridges the C-based Whisper.cpp library with Swift's modern concurrency model, handling everything from model initialization to memory-safe resource cleanup.

Core Functionalities of WhisperEngine.swift

The WhisperEngine class provides seven primary responsibilities that form a complete speech-to-text pipeline according to the OpenSuperWhisper source code.

Model Loading and Initialization

The initialize() method creates a Whisper context from a selected model file stored in AppPreferences. It constructs a WhisperContextParams object and instantiates the underlying C++ context via MyWhisperContext.initFromFile, preparing the engine for inference. This implementation resides at lines 65-77 of WhisperEngine.swift.

Audio Preparation and Conversion

Before inference, convertAudioToPCM(url:) converts any supplied audio to the 16kHz mono PCM format required by Whisper. The method handles MP4-to-M4A conversion when necessary, then selects between parallel or sequential processing paths based on file length. Using AVAudioConverter, it mixes multi-channel audio to mono while discarding silent channels through energy-based channel selection. This non-isolated function appears at lines 39-84.

Transcription Execution

The transcribeAudio(url:settings:) method configures WhisperFullParams—including beam-search settings, language detection, timestamps, and temperature—before invoking the C-level Whisper inference. It installs a C-level progress callback, calls context.full(samples:params:), then iterates over generated segments to optionally insert timestamps, remove placeholder tags, and apply Asian-language autocorrection via AutocorrectWrapper. This core logic spans lines 79-108 and lines 119-130.

Progress Reporting

Progress updates bridge Whisper's internal 0-100% range to a UI-friendly 5-95% scale via onProgressUpdate. The ProgressContext class stores the last reported value and forwards updates on the main queue, while the C-level progressCallback function bridges to Swift through an unmanaged opaque pointer. Implementation details appear at lines 5-22 and lines 36-51.

Cancellation Support

The cancelTranscription() method enables UI-driven abortion of running transcription jobs. It sets an atomic _isCancelled flag protected by stateLock (an NSLock instance), then flips the abortFlag pointer that Whisper checks through its C-level abort_callback. This thread-safe implementation occupies lines 110-115 and lines 28-58.

Language Support

getSupportedLanguages() exposes the set of languages Whisper can handle by delegating to LanguageUtil.availableLanguages, located at lines 117-120.

Utility Helpers

Supporting functions include resolveFileURL for file-type detection, makeTargetFormat for target audio format creation, and appendMixedSamples for channel mixing and energy-based channel selection. These utilities appear at lines 121-137, lines 192-203, and lines 226-242 respectively.

Architectural Design Patterns

Beyond functional implementation, WhisperEngine.swift employs specific architectural patterns to ensure safety and flexibility in the OpenSuperWhisper codebase.

Protocol-Driven Abstraction

WhisperEngine conforms to the TranscriptionEngine protocol defined in OpenSuperWhisper/Engines/TranscriptionEngine.swift, allowing OpenSuperWhisper to swap in alternative engines (such as FluidAudioEngine) without modifying downstream code.

Concurrency and Thread Safety

Audio conversion runs on a detached Task with .userInitiated priority, utilizing parallel workers for large files. The actual Whisper inference executes on the calling async context, while all UI callbacks marshal onto the main queue. Mutable state—including _isCancelled and _abortFlag—is protected by stateLock, ensuring thread-safe access across concurrent operations.

C-Swift Bridge Safety

Whisper's native C API requires raw function pointers for abort and progress callbacks. The engine safely bridges these by passing an unmanaged ProgressContext and a mutable Bool flag, ensuring memory safety and proper thread synchronization. All temporary resources are de-allocated in defer blocks to prevent leaks.

Implementation Examples

The following patterns demonstrate practical usage of WhisperEngine.swift in OpenSuperWhisper workflows.

Basic Initialization and Transcription

let engine = WhisperEngine()
engine.onProgressUpdate = { progress in
    print("Transcription progress:", progress)
}

Task {
    do {
        try await engine.initialize()
        let audioURL = Bundle.main.url(forResource: "sample", withExtension: "wav")!
        let settings = Settings()
        let text = try await engine.transcribeAudio(url: audioURL, settings: settings)
        print("Result:", text)
    } catch {
        print("Error:", error)
    }
}

Cancelling Long-Running Transcription

let engine = WhisperEngine()
Task {
    do {
        try await engine.initialize()
        let transcriptionTask = Task {
            try await engine.transcribeAudio(url: longAudioURL, settings: Settings())
        }
        
        // Cancel when user taps "Stop"
        engine.cancelTranscription()
        let _ = await transcriptionTask.result
    } catch {
        print("Cancelled or failed:", error)
    }
}

Querying Supported Languages

let engine = WhisperEngine()
let languages = engine.getSupportedLanguages()
print("Whisper supports:", languages)

Summary

  • WhisperEngine.swift serves as the concrete Whisper implementation in OpenSuperWhisper, located at OpenSuperWhisper/Engines/WhisperEngine.swift.
  • The engine handles the complete pipeline: model loading via initialize(), 16kHz mono PCM conversion via convertAudioToPCM(), and inference through transcribeAudio(url:settings:).
  • Thread-safe cancellation is implemented through NSLock-protected flags and C-level callback bridging.
  • Real-time progress reporting maps Whisper's internal 0-100% progress to a UI-friendly 5-95% range.
  • The class conforms to the TranscriptionEngine protocol, enabling swappable backend architectures.

Frequently Asked Questions

How does WhisperEngine.swift handle different audio formats?

The convertAudioToPCM(url:) method automatically handles format conversion through AVAudioConverter, converting MP4 files to M4A when necessary and resampling all input to 16kHz mono PCM. It employs parallel processing for large files and sequential conversion for smaller ones, mixing multi-channel audio to mono while discarding silent channels based on energy detection.

What concurrency model does WhisperEngine.swift use?

Audio conversion runs on a detached Task with .userInitiated priority to prevent blocking the main thread, while Whisper inference executes on the calling async context. All UI-facing callbacks dispatch to the main queue. Mutable state variables are protected by stateLock, an NSLock instance ensuring thread-safe cancellation and flag updates.

How does the engine communicate progress to the UI?

WhisperEngine.swift maps Whisper's internal 0-100% progress to a 5-95% range through the ProgressContext class and forwards updates via the onProgressUpdate closure on the main queue. A C-level progressCallback function bridges Whisper's native API to Swift by passing an unmanaged ProgressContext pointer, ensuring zero-overhead progress reporting without retain cycles.

Can transcription be cancelled mid-processing?

Yes. The cancelTranscription() method sets an atomic _isCancelled flag protected by NSLock and updates the abortFlag pointer that Whisper checks through its C-level abort_callback. This allows immediate termination of long-running inference without waiting for the current processing chunk to complete, with all resources cleaned up in a defer block.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →