How to Use the Whisper Engine for Speech-to-Text on macOS with OpenSuperWhisper

OpenSuperWhisper provides a native Swift wrapper around OpenAI's Whisper that loads models, converts audio to 16 kHz mono PCM, and exposes an async transcribeAudio method for high-quality offline transcription on macOS.

The OpenSuperWhisper repository offers a complete implementation of the Whisper engine for speech-to-text on macOS, wrapping the C++ Whisper library in Swift to provide a native transcription experience. The architecture centers on the WhisperEngine class, which conforms to the TranscriptionEngine protocol and handles model loading, audio preprocessing, and inference. This guide walks through the core components, configuration requirements, and implementation patterns needed to integrate speech recognition into your macOS applications.

Core Architecture of the Whisper Engine

The speech-to-text pipeline in OpenSuperWhisper splits responsibilities across several specialized classes to maintain clean separation between audio processing, model inference, and UI coordination.

WhisperEngine.swift - The Main Interface

The WhisperEngine class in OpenSuperWhisper/Engines/WhisperEngine.swift serves as the primary implementation of the transcription protocol. It manages the lifecycle of the native Whisper context and exposes the transcribeAudio(url:settings:) method that your application calls to process audio files. When initialized, the engine creates a MyWhisperContext instance via MyWhisperContext.initFromFile, loading the model binary specified in AppPreferences.shared.selectedWhisperModelPath.

TranscriptionService.swift - Orchestration Layer

Rather than interacting directly with the engine, most UI code should use the TranscriptionService singleton accessed via TranscriptionService.shared. This singleton, defined in OpenSuperWhisper/TranscriptionService.swift, manages engine lifecycle through its loadEngine() method and provides the high-level async API transcribeAudio(url:settings:). The service automatically handles engine selection based on user preferences and publishes transcription state through @Published properties that SwiftUI views can bind to.

MyWhisperContext and the C-API Bridge

The actual inference happens in MyWhisperContext, a thin Swift wrapper around the Whisper C-API found in OpenSuperWhisper/Whis/Whis.swift. This class owns the native Whisper state and exposes the full(samples:params:&cParams) method that runs the decoding loop. The context also provides helpers like fullNSegments for retrieving individual text segments and their timestamps after inference completes.

Preparing Your Model and Settings

Before running transcription, you must configure the model path and transcription parameters that control decoding behavior.

Configuring AppPreferences for Model Paths

The engine requires a valid Whisper model file (e.g., ggml-tiny.en.bin) on disk. Set the path via AppPreferences.shared.selectedWhisperModelPath before calling the transcription service. If this property is nil or points to a non-existent file, TranscriptionService throws TranscriptionError.contextInitializationFailed when attempting to load the engine.

// Set the model path before transcription
AppPreferences.shared.selectedWhisperModelPath = "/Users/you/Models/ggml-tiny.en.bin"

Customizing Transcription Parameters

The Settings struct in OpenSuperWhisper/Settings.swift encapsulates all user-controllable transcription options. Create an instance of this value type to specify language detection, sampling strategies, and output formatting.

Key parameters include:

  • useBeamSearch - Enables beam search decoding when true (alternative to greedy sampling)
  • beamSize - Number of beams to maintain during beam search (default: 5)
  • temperature - Sampling temperature for token generation (0.0 for deterministic output)
  • selectedLanguage - ISO-639-1 language code (e.g., "en") or "auto" for automatic detection
  • showTimestamps - When true, inserts [t0->t1] timestamps into the output text
  • suppressBlankAudio - Reduces hallucinations by suppressing blank audio segments
var settings = Settings()
settings.useBeamSearch = true
settings.beamSize = 5
settings.temperature = 0.0
settings.selectedLanguage = "en"
settings.showTimestamps = true

The Transcription Workflow

When TranscriptionService.shared.transcribeAudio(url:settings:) is invoked, the engine executes a three-stage pipeline: audio conversion, model inference, and result formatting.

Audio Conversion to 16 kHz Mono

The WhisperEngine.convertAudioToPCM method preprocesses all input audio to meet Whisper's requirements. Using AVAudioFile, the engine reads the source audio and converts it to a 16 kHz mono format via makeTargetFormat(channelCount:). The method handles MP4 containers by copying non-native .m4a files to temporary locations before processing. For short clips, it uses a fast sequential converter (convertSequential), while longer files utilize a parallel conversion pipeline to optimize throughput. The resulting Float array represents the PCM samples passed to the native context.

Running Inference with WhisperFullParams

With PCM samples ready, the engine constructs a WhisperFullParams struct (defined in OpenSuperWhisper/Whis/WhisperFullParams.swift) populated from your Settings instance. This struct configures the decoding strategy, language constraints, and audio thresholds. The engine then calls context.full(samples:params:&cParams), which executes the native Whisper loop. During inference, the engine maps Whisper's native 0-100% progress to a UI-friendly 10-95% range through the progressCallback closure.

After completion, the engine iterates over context.fullNSegments, extracting text and timestamps to build the final transcription string.

Handling Progress and Cancellation

Both audio conversion and model inference respect a shared abortFlag (an UnsafeMutablePointer<Bool>). To cancel an ongoing transcription, call TranscriptionService.shared.cancelTranscription(), which flips the flag that the native Whisper loop checks through its abort_callback. Progress updates flow through the onProgressUpdate closure attached to the engine, allowing the UI to display real-time completion percentages.

Complete Implementation Example

The following self-contained Swift snippet demonstrates loading a model, configuring transcription settings, and executing speech-to-text on a local audio file:

import OpenSuperWhisper

// 1. Configure the model path
AppPreferences.shared.selectedWhisperModelPath = "/Users/you/Models/ggml-tiny.en.bin"

// 2. Set transcription parameters
var whisperSettings = Settings()
whisperSettings.useBeamSearch = true
whisperSettings.beamSize = 5
whisperSettings.temperature = 0.0
whisperSettings.selectedLanguage = "en"
whisperSettings.showTimestamps = true

// 3. Specify the audio file
let audioURL = URL(fileURLWithPath: "/Users/you/AudioSamples/interview.wav")

// 4. Execute transcription asynchronously
Task {
    do {
        let transcript = try await TranscriptionService.shared.transcribeAudio(
            url: audioURL,
            settings: whisperSettings
        )
        print("Transcription complete:\n\(transcript)")
    } catch {
        print("Transcription failed: \(error)")
    }
}

Summary

  • Model Requirement: Set AppPreferences.shared.selectedWhisperModelPath to a valid Whisper binary (e.g., ggml-tiny.en.bin) before transcribing.
  • Audio Processing: The WhisperEngine automatically converts input audio to 16 kHz mono PCM using AVAudioFile and parallel converters for large files.
  • Configuration: Use the Settings struct to control beam search, temperature, language detection, and timestamp formatting.
  • Threading: All heavy computation runs off the main thread via Task.detached, ensuring responsive macOS UIs.
  • Cancellation: Call TranscriptionService.shared.cancelTranscription() to abort processing via the shared abortFlag checked by the native C-API.

Frequently Asked Questions

What audio formats does the WhisperEngine support?

The engine uses AVAudioFile to read input, supporting any format Core Audio handles natively, including WAV, MP4, and M4A. The convertAudioToPCM method in WhisperEngine.swift specifically handles MP4 container edge cases by copying non-native .m4a files to temporary locations before conversion, ensuring robust format support.

How do I cancel an ongoing transcription?

Call TranscriptionService.shared.cancelTranscription() to immediately halt processing. This method sets the shared abortFlag that both the audio conversion pipeline and the native Whisper inference loop check periodically via the abort_callback mechanism, allowing safe termination without blocking the main thread.

Why does the engine require 16 kHz mono audio?

Whisper models are trained on 16 kHz mono PCM data. The makeTargetFormat(channelCount:) method in WhisperEngine.swift ensures compliance by resampling input audio and downmixing stereo channels to mono. Supplying audio in this format prevents frequency aliasing and ensures the model processes the correct spectral information.

Where should I store the Whisper model files?

Store model binaries (e.g., ggml-tiny.en.bin, ggml-base.en.bin) anywhere in the filesystem accessible to your app, then reference the absolute path via AppPreferences.shared.selectedWhisperModelPath. In production applications, consider bundling models in the app bundle or downloading them to Application Support directories, ensuring the path is set before the first call to loadEngine().

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →