# How to Use OpenSuperWhisper for Audio Transcription: A Complete Guide

> Learn how to use OpenSuperWhisper for audio transcription on macOS. This guide covers real-time capture, format conversion, and on-device transcription with Whisper or FluidAudio.

- Repository: [Starmel/OpenSuperWhisper](https://github.com/Starmel/OpenSuperWhisper)
- Tags: how-to-guide
- Published: 2026-07-07

---

**OpenSuperWhisper is a macOS-native application that captures real-time microphone input, converts audio to Whisper-compatible PCM format, and performs on-device transcription using Whisper or FluidAudio engines.**

OpenSuperWhisper, hosted in the Starmel/OpenSuperWhisper repository, provides a complete Swift implementation for local audio transcription using OpenAI's Whisper models. The application processes audio entirely on-device through a three-stage pipeline involving capture, conversion, and transcription, exposing a clean API for integration into macOS workflows.

## Three-Stage Transcription Pipeline

The transcription workflow follows a structured pipeline implemented across [`AudioRecorder.swift`](https://github.com/Starmel/OpenSuperWhisper/blob/main/AudioRecorder.swift) and [`WhisperEngine.swift`](https://github.com/Starmel/OpenSuperWhisper/blob/main/WhisperEngine.swift).

### Audio Capture

The `AudioRecorder` class manages microphone access using a private serial queue (`workQueue`) to prevent UI blocking. It writes 16-bit PCM audio at 16 kHz using `AVFormatIDKey`, `AVSampleRateKey`, and `AVLinearPCMBitDepthKey` configurations.

Key features include:
- Temporary `.wav` file creation for each recording session
- Automatic cleanup of old temporary files via `cleanupOldTemporaryFiles()`
- Optional start-notification sounds for audio feedback
- Support for global hotkeys, modifier keys, or mouse button triggers

### Audio Conversion

Before transcription, `WhisperEngine.convertAudioToPCM` processes recorded or user-provided audio into the format required by Whisper models. The conversion implementation:

- Resamples audio to 16 kHz mono Float32 samples using `AVAudioConverter`
- Builds target formats via `makeTargetFormat`
- Parallelizes conversion across multiple cores for large files

This ensures compatibility regardless of source audio format or sample rate.

### Transcription and Post-Processing

The `WhisperEngine.transcribeAudio` method loads the model from `AppPreferences.shared.selectedWhisperModelPath` and initializes a fresh decoding state for each session. The transcription pipeline includes:

- **Voice Activity Detection**: Uses the built-in Silero VAD model (`ggml-silero-v5.1.2.bin`) to remove silence and prevent hallucinations
- **Progress Reporting**: Maps Whisper's 0-100% callbacks to the app's 10-95% range via `ProgressContext`
- **Token Filtering**: Removes placeholder tokens including `[MUSIC]` and `[BLANK_AUDIO]`
- **Language-Specific Formatting**: Applies `AutocorrectWrapper.format` for Asian language autocorrection when enabled

The UI layer in [`ContentView.swift`](https://github.com/Starmel/OpenSuperWhisper/blob/main/ContentView.swift) binds to `AudioRecorder.isRecording`, `AudioRecorder.isPlaying`, and `TranscriptionEngine` progress callbacks to display live indicators and final transcripts.

## Implementation Architecture

Understanding the internal structure enables proper integration and extension of OpenSuperWhisper functionality.

### Recording System

`AudioRecorder` operates on a dedicated serial dispatch queue to maintain responsive UI during capture. The class handles microphone permissions, temporary file management, and playback functionality. Recordings persist as temporary `.wav` files until transcription completes or cleanup routines execute.

### Model Management

[`WhisperModelManager.swift`](https://github.com/Starmel/OpenSuperWhisper/blob/main/WhisperModelManager.swift) handles downloading and selecting Whisper model files. The engine initializes via `WhisperEngine.initialize()`, which loads the model specified in `AppPreferences.shared.selectedWhisperModelPath`. The `TranscriptionEngine` protocol abstraction allows seamless switching between `WhisperEngine` and `FluidAudioEngine` implementations.

### Configuration Interface

[`Settings.swift`](https://github.com/Starmel/OpenSuperWhisper/blob/main/Settings.swift) stores user-configurable transcription parameters including language selection, beam search width, and temperature controls. These settings pass directly to the transcription engine via the `Settings.shared` singleton.

## Code Examples

### Basic Recording and Transcription Workflow

The following Swift code demonstrates the complete workflow from hotkey-triggered recording to final transcript:

```swift
// Start recording (triggered by global hotkey, modifier key, or mouse button)
AudioRecorder.shared.startRecording()

// Stop recording and transcribe
if let wavURL = await AudioRecorder.shared.stopRecording() {
    let engine: TranscriptionEngine = WhisperEngine()  // or FluidAudioEngine()
    try await engine.initialize()
    
    let transcript = try await engine.transcribeAudio(
        url: wavURL,
        settings: Settings.shared  // Contains language, beam search, etc.
    )
    print("Transcript: \(transcript)")
}

```

### Manual Audio Conversion

For testing or batch processing existing audio files:

```swift
let engine = WhisperEngine()
try await engine.initialize()

if let pcmSamples = try await engine.convertAudioToPCM(fileURL: someAudioURL) {
    print("Converted to \(pcmSamples.count) Float32 samples")
}

```

### Querying Supported Languages

Populate UI language pickers using the engine's capabilities:

```swift
let supported = WhisperEngine().getSupportedLanguages()
print("Supported languages: \(supported.joined(separator: ", "))")

```

### Integrating Progress Callbacks

Monitor transcription progress for UI updates:

```swift
// Progress reported via ProgressContext
// Maps Whisper's internal 0-100% to application's 10-95% range
// Bind to UI progress indicators in ContentView.swift

```

## Configuration Options

OpenSuperWhisper exposes several customization points through [`Settings.swift`](https://github.com/Starmel/OpenSuperWhisper/blob/main/Settings.swift):

- **Language Selection**: Specify source language or enable auto-detection
- **Beam Search Parameters**: Configure beam width for decoding accuracy
- **Temperature**: Control randomness in token generation
- **VAD Sensitivity**: Adjust Silero VAD thresholds to filter silence

The [`ContentView.swift`](https://github.com/Starmel/OpenSuperWhisper/blob/main/ContentView.swift) interface supports drag-and-drop functionality for queuing additional audio files after initial transcription.

## Summary

- OpenSuperWhisper implements a complete on-device transcription pipeline using [`AudioRecorder.swift`](https://github.com/Starmel/OpenSuperWhisper/blob/main/AudioRecorder.swift) for capture and [`WhisperEngine.swift`](https://github.com/Starmel/OpenSuperWhisper/blob/main/WhisperEngine.swift) for processing
- Audio converts to 16 kHz mono Float32 PCM format before transcription, with optional multicore parallelization for large files
- Voice Activity Detection using the Silero model (`ggml-silero-v5.1.2.bin`) filters silence to reduce hallucinations and improve accuracy
- The `TranscriptionEngine` protocol supports both Whisper and FluidAudio backends through a unified interface
- Progress reporting maps internal Whisper callbacks to the 10-95% UI range via `ProgressContext` for accurate progress indicators

## Frequently Asked Questions

### How does OpenSuperWhisper handle audio format conversion?

According to the Starmel/OpenSuperWhisper source code in [`WhisperEngine.swift`](https://github.com/Starmel/OpenSuperWhisper/blob/main/WhisperEngine.swift), the `convertAudioToPCM` method uses `AVAudioConverter` with a target format built by `makeTargetFormat` to resample any input audio to 16 kHz mono Float32 samples. For large files, the conversion can run in parallel across multiple CPU cores to improve performance.

### What is the purpose of the Silero VAD model in OpenSuperWhisper?

The `ggml-silero-v5.1.2.bin` model performs Voice Activity Detection before transcription begins. As implemented in [`WhisperEngine.swift`](https://github.com/Starmel/OpenSuperWhisper/blob/main/WhisperEngine.swift), this preprocessing step identifies and removes silent segments from the audio stream using the `context.full(samples:params:)` method, significantly reducing transcription hallucinations and improving processing speed by skipping non-speech content.

### Can I use OpenSuperWhisper with pre-recorded audio files?

Yes. While the default workflow captures microphone input via `AudioRecorder.startRecording()`, you can pass any audio file URL directly to `WhisperEngine.transcribeAudio` or manually convert files using `convertAudioToPCM`. The UI in [`ContentView.swift`](https://github.com/Starmel/OpenSuperWhisper/blob/main/ContentView.swift) supports drag-and-drop functionality for queuing existing audio files, and the system processes them through the same conversion and transcription pipeline.

### How does the application prevent UI freezing during recording?

[`AudioRecorder.swift`](https://github.com/Starmel/OpenSuperWhisper/blob/main/AudioRecorder.swift) runs all recording operations on a private serial queue (`workQueue`) rather than the main thread. This architecture ensures that real-time audio capture, file I/O operations, and `cleanupOldTemporaryFiles` execution do not block the SwiftUI interface, allowing smooth updates to recording indicators via `AudioRecorder.isRecording` and `AudioRecorder.isPlaying` bindings.