# How to Use the Whisper Engine for Speech-to-Text on macOS with OpenSuperWhisper

> Learn to use the Whisper engine for speech to text on macOS with OpenSuperWhisper. Get high-quality offline transcription using this native Swift wrapper and its async transcribeAudio method.

- Repository: [Starmel/OpenSuperWhisper](https://github.com/Starmel/OpenSuperWhisper)
- Tags: how-to-guide
- Published: 2026-07-05

---

**OpenSuperWhisper provides a native Swift wrapper around OpenAI's Whisper that loads models, converts audio to 16 kHz mono PCM, and exposes an async `transcribeAudio` method for high-quality offline transcription on macOS.**

The OpenSuperWhisper repository offers a complete implementation of the **Whisper engine for speech-to-text on macOS**, wrapping the C++ Whisper library in Swift to provide a native transcription experience. The architecture centers on the `WhisperEngine` class, which conforms to the `TranscriptionEngine` protocol and handles model loading, audio preprocessing, and inference. This guide walks through the core components, configuration requirements, and implementation patterns needed to integrate speech recognition into your macOS applications.

## Core Architecture of the Whisper Engine

The speech-to-text pipeline in OpenSuperWhisper splits responsibilities across several specialized classes to maintain clean separation between audio processing, model inference, and UI coordination.

### WhisperEngine.swift - The Main Interface

The `WhisperEngine` class in [`OpenSuperWhisper/Engines/WhisperEngine.swift`](https://github.com/Starmel/OpenSuperWhisper/blob/main/OpenSuperWhisper/Engines/WhisperEngine.swift) serves as the primary implementation of the transcription protocol. It manages the lifecycle of the native Whisper context and exposes the `transcribeAudio(url:settings:)` method that your application calls to process audio files. When initialized, the engine creates a `MyWhisperContext` instance via `MyWhisperContext.initFromFile`, loading the model binary specified in `AppPreferences.shared.selectedWhisperModelPath`.

### TranscriptionService.swift - Orchestration Layer

Rather than interacting directly with the engine, most UI code should use the `TranscriptionService` singleton accessed via `TranscriptionService.shared`. This singleton, defined in [`OpenSuperWhisper/TranscriptionService.swift`](https://github.com/Starmel/OpenSuperWhisper/blob/main/OpenSuperWhisper/TranscriptionService.swift), manages engine lifecycle through its `loadEngine()` method and provides the high-level async API `transcribeAudio(url:settings:)`. The service automatically handles engine selection based on user preferences and publishes transcription state through `@Published` properties that SwiftUI views can bind to.

### MyWhisperContext and the C-API Bridge

The actual inference happens in `MyWhisperContext`, a thin Swift wrapper around the Whisper C-API found in [`OpenSuperWhisper/Whis/Whis.swift`](https://github.com/Starmel/OpenSuperWhisper/blob/main/OpenSuperWhisper/Whis/Whis.swift). This class owns the native Whisper state and exposes the `full(samples:params:&cParams)` method that runs the decoding loop. The context also provides helpers like `fullNSegments` for retrieving individual text segments and their timestamps after inference completes.

## Preparing Your Model and Settings

Before running transcription, you must configure the model path and transcription parameters that control decoding behavior.

### Configuring AppPreferences for Model Paths

The engine requires a valid Whisper model file (e.g., `ggml-tiny.en.bin`) on disk. Set the path via `AppPreferences.shared.selectedWhisperModelPath` before calling the transcription service. If this property is `nil` or points to a non-existent file, `TranscriptionService` throws `TranscriptionError.contextInitializationFailed` when attempting to load the engine.

```swift
// Set the model path before transcription
AppPreferences.shared.selectedWhisperModelPath = "/Users/you/Models/ggml-tiny.en.bin"

```

### Customizing Transcription Parameters

The `Settings` struct in [`OpenSuperWhisper/Settings.swift`](https://github.com/Starmel/OpenSuperWhisper/blob/main/OpenSuperWhisper/Settings.swift) encapsulates all user-controllable transcription options. Create an instance of this value type to specify language detection, sampling strategies, and output formatting.

Key parameters include:

- **`useBeamSearch`** - Enables beam search decoding when `true` (alternative to greedy sampling)
- **`beamSize`** - Number of beams to maintain during beam search (default: 5)
- **`temperature`** - Sampling temperature for token generation (0.0 for deterministic output)
- **`selectedLanguage`** - ISO-639-1 language code (e.g., "en") or "auto" for automatic detection
- **`showTimestamps`** - When `true`, inserts `[t0->t1]` timestamps into the output text
- **`suppressBlankAudio`** - Reduces hallucinations by suppressing blank audio segments

```swift
var settings = Settings()
settings.useBeamSearch = true
settings.beamSize = 5
settings.temperature = 0.0
settings.selectedLanguage = "en"
settings.showTimestamps = true

```

## The Transcription Workflow

When `TranscriptionService.shared.transcribeAudio(url:settings:)` is invoked, the engine executes a three-stage pipeline: audio conversion, model inference, and result formatting.

### Audio Conversion to 16 kHz Mono

The `WhisperEngine.convertAudioToPCM` method preprocesses all input audio to meet Whisper's requirements. Using `AVAudioFile`, the engine reads the source audio and converts it to a 16 kHz mono format via `makeTargetFormat(channelCount:)`. The method handles MP4 containers by copying non-native `.m4a` files to temporary locations before processing. For short clips, it uses a fast sequential converter (`convertSequential`), while longer files utilize a parallel conversion pipeline to optimize throughput. The resulting `Float` array represents the PCM samples passed to the native context.

### Running Inference with WhisperFullParams

With PCM samples ready, the engine constructs a `WhisperFullParams` struct (defined in [`OpenSuperWhisper/Whis/WhisperFullParams.swift`](https://github.com/Starmel/OpenSuperWhisper/blob/main/OpenSuperWhisper/Whis/WhisperFullParams.swift)) populated from your `Settings` instance. This struct configures the decoding strategy, language constraints, and audio thresholds. The engine then calls `context.full(samples:params:&cParams)`, which executes the native Whisper loop. During inference, the engine maps Whisper's native 0-100% progress to a UI-friendly 10-95% range through the `progressCallback` closure.

After completion, the engine iterates over `context.fullNSegments`, extracting text and timestamps to build the final transcription string.

### Handling Progress and Cancellation

Both audio conversion and model inference respect a shared `abortFlag` (an `UnsafeMutablePointer<Bool>`). To cancel an ongoing transcription, call `TranscriptionService.shared.cancelTranscription()`, which flips the flag that the native Whisper loop checks through its `abort_callback`. Progress updates flow through the `onProgressUpdate` closure attached to the engine, allowing the UI to display real-time completion percentages.

## Complete Implementation Example

The following self-contained Swift snippet demonstrates loading a model, configuring transcription settings, and executing speech-to-text on a local audio file:

```swift
import OpenSuperWhisper

// 1. Configure the model path
AppPreferences.shared.selectedWhisperModelPath = "/Users/you/Models/ggml-tiny.en.bin"

// 2. Set transcription parameters
var whisperSettings = Settings()
whisperSettings.useBeamSearch = true
whisperSettings.beamSize = 5
whisperSettings.temperature = 0.0
whisperSettings.selectedLanguage = "en"
whisperSettings.showTimestamps = true

// 3. Specify the audio file
let audioURL = URL(fileURLWithPath: "/Users/you/AudioSamples/interview.wav")

// 4. Execute transcription asynchronously
Task {
    do {
        let transcript = try await TranscriptionService.shared.transcribeAudio(
            url: audioURL,
            settings: whisperSettings
        )
        print("Transcription complete:\n\(transcript)")
    } catch {
        print("Transcription failed: \(error)")
    }
}

```

## Summary

- **Model Requirement**: Set `AppPreferences.shared.selectedWhisperModelPath` to a valid Whisper binary (e.g., `ggml-tiny.en.bin`) before transcribing.
- **Audio Processing**: The `WhisperEngine` automatically converts input audio to 16 kHz mono PCM using `AVAudioFile` and parallel converters for large files.
- **Configuration**: Use the `Settings` struct to control beam search, temperature, language detection, and timestamp formatting.
- **Threading**: All heavy computation runs off the main thread via `Task.detached`, ensuring responsive macOS UIs.
- **Cancellation**: Call `TranscriptionService.shared.cancelTranscription()` to abort processing via the shared `abortFlag` checked by the native C-API.

## Frequently Asked Questions

### What audio formats does the WhisperEngine support?

The engine uses `AVAudioFile` to read input, supporting any format Core Audio handles natively, including WAV, MP4, and M4A. The `convertAudioToPCM` method in [`WhisperEngine.swift`](https://github.com/Starmel/OpenSuperWhisper/blob/main/WhisperEngine.swift) specifically handles MP4 container edge cases by copying non-native `.m4a` files to temporary locations before conversion, ensuring robust format support.

### How do I cancel an ongoing transcription?

Call `TranscriptionService.shared.cancelTranscription()` to immediately halt processing. This method sets the shared `abortFlag` that both the audio conversion pipeline and the native Whisper inference loop check periodically via the `abort_callback` mechanism, allowing safe termination without blocking the main thread.

### Why does the engine require 16 kHz mono audio?

Whisper models are trained on 16 kHz mono PCM data. The `makeTargetFormat(channelCount:)` method in [`WhisperEngine.swift`](https://github.com/Starmel/OpenSuperWhisper/blob/main/WhisperEngine.swift) ensures compliance by resampling input audio and downmixing stereo channels to mono. Supplying audio in this format prevents frequency aliasing and ensures the model processes the correct spectral information.

### Where should I store the Whisper model files?

Store model binaries (e.g., `ggml-tiny.en.bin`, `ggml-base.en.bin`) anywhere in the filesystem accessible to your app, then reference the absolute path via `AppPreferences.shared.selectedWhisperModelPath`. In production applications, consider bundling models in the app bundle or downloading them to Application Support directories, ensuring the path is set before the first call to `loadEngine()`.