How to Use OpenSuperWhisper for Audio Transcription: A Complete Guide

OpenSuperWhisper is a macOS-native application that captures real-time microphone input, converts audio to Whisper-compatible PCM format, and performs on-device transcription using Whisper or FluidAudio engines.

OpenSuperWhisper, hosted in the Starmel/OpenSuperWhisper repository, provides a complete Swift implementation for local audio transcription using OpenAI's Whisper models. The application processes audio entirely on-device through a three-stage pipeline involving capture, conversion, and transcription, exposing a clean API for integration into macOS workflows.

Three-Stage Transcription Pipeline

The transcription workflow follows a structured pipeline implemented across AudioRecorder.swift and WhisperEngine.swift.

Audio Capture

The AudioRecorder class manages microphone access using a private serial queue (workQueue) to prevent UI blocking. It writes 16-bit PCM audio at 16 kHz using AVFormatIDKey, AVSampleRateKey, and AVLinearPCMBitDepthKey configurations.

Key features include:

  • Temporary .wav file creation for each recording session
  • Automatic cleanup of old temporary files via cleanupOldTemporaryFiles()
  • Optional start-notification sounds for audio feedback
  • Support for global hotkeys, modifier keys, or mouse button triggers

Audio Conversion

Before transcription, WhisperEngine.convertAudioToPCM processes recorded or user-provided audio into the format required by Whisper models. The conversion implementation:

  • Resamples audio to 16 kHz mono Float32 samples using AVAudioConverter
  • Builds target formats via makeTargetFormat
  • Parallelizes conversion across multiple cores for large files

This ensures compatibility regardless of source audio format or sample rate.

Transcription and Post-Processing

The WhisperEngine.transcribeAudio method loads the model from AppPreferences.shared.selectedWhisperModelPath and initializes a fresh decoding state for each session. The transcription pipeline includes:

  • Voice Activity Detection: Uses the built-in Silero VAD model (ggml-silero-v5.1.2.bin) to remove silence and prevent hallucinations
  • Progress Reporting: Maps Whisper's 0-100% callbacks to the app's 10-95% range via ProgressContext
  • Token Filtering: Removes placeholder tokens including [MUSIC] and [BLANK_AUDIO]
  • Language-Specific Formatting: Applies AutocorrectWrapper.format for Asian language autocorrection when enabled

The UI layer in ContentView.swift binds to AudioRecorder.isRecording, AudioRecorder.isPlaying, and TranscriptionEngine progress callbacks to display live indicators and final transcripts.

Implementation Architecture

Understanding the internal structure enables proper integration and extension of OpenSuperWhisper functionality.

Recording System

AudioRecorder operates on a dedicated serial dispatch queue to maintain responsive UI during capture. The class handles microphone permissions, temporary file management, and playback functionality. Recordings persist as temporary .wav files until transcription completes or cleanup routines execute.

Model Management

WhisperModelManager.swift handles downloading and selecting Whisper model files. The engine initializes via WhisperEngine.initialize(), which loads the model specified in AppPreferences.shared.selectedWhisperModelPath. The TranscriptionEngine protocol abstraction allows seamless switching between WhisperEngine and FluidAudioEngine implementations.

Configuration Interface

Settings.swift stores user-configurable transcription parameters including language selection, beam search width, and temperature controls. These settings pass directly to the transcription engine via the Settings.shared singleton.

Code Examples

Basic Recording and Transcription Workflow

The following Swift code demonstrates the complete workflow from hotkey-triggered recording to final transcript:

// Start recording (triggered by global hotkey, modifier key, or mouse button)
AudioRecorder.shared.startRecording()

// Stop recording and transcribe
if let wavURL = await AudioRecorder.shared.stopRecording() {
    let engine: TranscriptionEngine = WhisperEngine()  // or FluidAudioEngine()
    try await engine.initialize()
    
    let transcript = try await engine.transcribeAudio(
        url: wavURL,
        settings: Settings.shared  // Contains language, beam search, etc.
    )
    print("Transcript: \(transcript)")
}

Manual Audio Conversion

For testing or batch processing existing audio files:

let engine = WhisperEngine()
try await engine.initialize()

if let pcmSamples = try await engine.convertAudioToPCM(fileURL: someAudioURL) {
    print("Converted to \(pcmSamples.count) Float32 samples")
}

Querying Supported Languages

Populate UI language pickers using the engine's capabilities:

let supported = WhisperEngine().getSupportedLanguages()
print("Supported languages: \(supported.joined(separator: ", "))")

Integrating Progress Callbacks

Monitor transcription progress for UI updates:

// Progress reported via ProgressContext
// Maps Whisper's internal 0-100% to application's 10-95% range
// Bind to UI progress indicators in ContentView.swift

Configuration Options

OpenSuperWhisper exposes several customization points through Settings.swift:

  • Language Selection: Specify source language or enable auto-detection
  • Beam Search Parameters: Configure beam width for decoding accuracy
  • Temperature: Control randomness in token generation
  • VAD Sensitivity: Adjust Silero VAD thresholds to filter silence

The ContentView.swift interface supports drag-and-drop functionality for queuing additional audio files after initial transcription.

Summary

  • OpenSuperWhisper implements a complete on-device transcription pipeline using AudioRecorder.swift for capture and WhisperEngine.swift for processing
  • Audio converts to 16 kHz mono Float32 PCM format before transcription, with optional multicore parallelization for large files
  • Voice Activity Detection using the Silero model (ggml-silero-v5.1.2.bin) filters silence to reduce hallucinations and improve accuracy
  • The TranscriptionEngine protocol supports both Whisper and FluidAudio backends through a unified interface
  • Progress reporting maps internal Whisper callbacks to the 10-95% UI range via ProgressContext for accurate progress indicators

Frequently Asked Questions

How does OpenSuperWhisper handle audio format conversion?

According to the Starmel/OpenSuperWhisper source code in WhisperEngine.swift, the convertAudioToPCM method uses AVAudioConverter with a target format built by makeTargetFormat to resample any input audio to 16 kHz mono Float32 samples. For large files, the conversion can run in parallel across multiple CPU cores to improve performance.

What is the purpose of the Silero VAD model in OpenSuperWhisper?

The ggml-silero-v5.1.2.bin model performs Voice Activity Detection before transcription begins. As implemented in WhisperEngine.swift, this preprocessing step identifies and removes silent segments from the audio stream using the context.full(samples:params:) method, significantly reducing transcription hallucinations and improving processing speed by skipping non-speech content.

Can I use OpenSuperWhisper with pre-recorded audio files?

Yes. While the default workflow captures microphone input via AudioRecorder.startRecording(), you can pass any audio file URL directly to WhisperEngine.transcribeAudio or manually convert files using convertAudioToPCM. The UI in ContentView.swift supports drag-and-drop functionality for queuing existing audio files, and the system processes them through the same conversion and transcription pipeline.

How does the application prevent UI freezing during recording?

AudioRecorder.swift runs all recording operations on a private serial queue (workQueue) rather than the main thread. This architecture ensures that real-time audio capture, file I/O operations, and cleanupOldTemporaryFiles execution do not block the SwiftUI interface, allowing smooth updates to recording indicators via AudioRecorder.isRecording and AudioRecorder.isPlaying bindings.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →